An automatic tracking system based on XR space
By establishing a preset instruction library and virtual scene division based on eye movement data in the XR system, the problem that traditional XR systems cannot adjust user operation feedback in real time is solved, achieving more natural eye movement interaction and higher user experience.
Patent Information
- Application Number
- CN202510228079.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Traditional XR systems lack real-time monitoring and adjustment mechanisms for user operation feedback, which affects users' immersion and operation experience, and cannot make dynamic adjustments based on users' vision and feelings.
By establishing a preset instruction library, rough and fine virtual scene division is performed based on eye movement data, the virtual scene where the user is located is judged in real time, the initial rotation value and boot time are set, the preliminary calculation model is dynamically adjusted, and the boot simulation is optimized.
It realizes a more natural eye movement interaction, improves user satisfaction and experience, and enhances the adaptability of virtual scenes and user immersion.
Smart Images

Figure CN119722746B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of XR space, and in particular to an automatic tracking system based on XR space. Background Art
[0002] Extended Reality (XR) refers to the application and implementation of Extended Reality (XR) technology in a specific environment. Its technology integrates virtual reality (VR), augmented reality (AR) and mixed reality (MR) to provide users with an immersive experience. In XR systems, the basic principle of eye tracking technology is similar to traditional eye tracking, which is to infer the user's gaze point and eye movement trajectory by monitoring the user's eye movement in real time. It is usually integrated with head-mounted devices (such as VR helmets, AR glasses, etc.) and relies on infrared cameras and optical sensors to capture the user's eye movement and pupil position.
[0003] However, traditional XR systems lack a real-time monitoring and adjustment mechanism for user operation feedback. For example, when a user is experiencing a game in an XR space, if the user repeatedly stays on the same task interface, the traditional solution usually relies on a handle or external force to guide the user. This is not only time-consuming and labor-intensive, but also unable to be dynamically adjusted according to the user's line of sight and feelings, affecting the user's sense of immersion and operating experience, and resulting in an inability to quickly and effectively solve the current problem. When the user looks at a specific area, the display brightness, contrast and other settings cannot be adjusted according to the eye movement characteristics of each user. To a certain extent, the virtual world may feel inconsistent with the real world, reducing the accurate identification of the user's position. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] In view of the shortcomings of the prior art, the present invention provides an automatic tracking system based on XR space, which obtains the gaze area and movement direction in the XR space by establishing a preset instruction library; performs coarse and fine virtual scene division based on eye movement data, judges the virtual scene where the user is in real time, sets the initial rotation value and guidance time, and dynamically adjusts the preliminary calculation model according to user feedback, optimizes the guidance simulation, ensures that the user obtains a realistic experience, achieves a more natural eye movement interaction, and integrates more deeply into the XR environment, improves user satisfaction and experience, and solves the problems raised in the background technology.
[0006] (II) Technical solution
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0008] An automatic tracking system based on XR space, including a gaze tracking module, a scene matching module, an interactive navigation module and a feedback adjustment module;
[0009] The eye tracking module identifies the eye position based on the real-time collected eye images, maps the eye position to the corresponding analysis instructions in the XR space based on the pre-built analysis instruction library, and identifies the movement direction and gaze area of the eye.
[0010] The scene matching module obtains a first virtual scene by roughly dividing the scene based on the movement direction and the gaze area; obtains a second virtual scene by finely dividing the first virtual scene by introducing an attention mechanism, and establishes an association relationship model based on the association relationship between the eye movement data and the corresponding virtual scene, and the association relationship includes a first association relationship and a second association relationship; wherein the eye movement data at least includes a movement direction and a gaze area;
[0011] The interactive navigation module guides users to interact in the XR space through the association model;
[0012] The feedback adjustment module identifies the position of the virtual scene where the current user is located, and builds a preliminary calculation model based on the initial rotation value and the guidance time under the condition that the user interacts according to the guidance, and generates a preset rotation value; then the feedback data of the user is counted, and the preliminary calculation model is adjusted through analysis and processing to perform recalibration of the user's position.
[0013] Furthermore, the step of identifying the movement direction and gaze area of the eyeball includes:
[0014] Image recognition: by collecting eye images, identifying the frame difference images of the eyes, performing differential processing on the frame difference images, obtaining the first two frame difference images and the last two frame difference images, and performing AND operations to obtain the phase difference images;
[0015] State recognition: Based on phase and difference images, computer vision and machine learning algorithms are used to capture the movement trajectory points of the eyeball. The movement trajectory points are the normalized sequence of the eyeball rotation trajectory points.
[0016] Motion mapping: Based on the motion trajectory points and pupil position, the long axis radius and short axis radius are identified, and the eye posture is calculated in combination with the environmental impact coefficient. This is matched with the predefined analysis instructions to identify the eye movement direction and gaze area and output them.
[0017] The steps for obtaining the environmental impact coefficient are as follows: collecting light, temperature, humidity, flatness and network bandwidth based on the motion trajectory points, and performing standardization processing according to the formula:
[0018] ;
[0019] Where enc represents the environmental impact coefficient, p1, p2, p3, p4 and p5 are the standardized light, temperature, humidity, flatness and network bandwidth. , , , as well as are weight correction coefficients, , as well as is a constant affecting parameter.
[0020] Furthermore, the steps of determining the virtual environment in which the user is located are:
[0021] Rough recognition: Divide the XR space into several different fields of view based on the gaze area and movement direction. In each field of view, mark the user's gaze time and movement direction to generate high and low priority corresponding focus areas; load, unload or update the scene content in each focus area, and obtain the first virtual scene; the focus area includes the foreground area, the background area and the interactive area;
[0022] Fine recognition: Based on the first virtual scene, an attention mechanism is introduced to dynamically enhance high-priority attention areas and dynamically weaken low-priority attention areas, and the scene content is reconstructed to update the first virtual scene to obtain the second virtual scene.
[0023] Furthermore, the steps of establishing the association relationship between the eye movement data and the virtual scene are:
[0024] Primary association: extracting the user's eye movement data based on the first virtual scene, creating a first time tag for the first virtual scene according to a time sequence, and establishing a first association relationship between the eye movement data and the first virtual scene;
[0025] Secondary association: based on the second virtual scene, extracting the eye movement data again, creating a second time tag for the second virtual scene according to the time sequence, and establishing a second association relationship between the eye movement data and the second virtual scene;
[0026] Model building: Based on the first association relationship, the second association relationship, the first virtual scene, the second virtual scene and the corresponding eye movement data, a relationship model between the user and the virtual scene is built.
[0027] Furthermore, the steps of guiding the user to interact in the XR space through the association relationship model are as follows:
[0028] Feature extraction: feature extraction technology is used to extract target features in the association model and construct a four-dimensional feature vector, including task, semantics, behavior, and time, which are marked as H, S, B, and time respectively;
[0029] In-depth analysis: Preset the time period T. For any time period Ta, construct several different combinations of judgment vectors through feature vectors to calculate the correlation between tasks, semantics and behaviors, evaluate the importance of task-semantics, task-behavior and semantics-behavior, and set the corresponding priorities according to the importance to perform interactive navigation in sequence; among them, the correlation is calculated based on the cosine similarity, and the priority is calculated as follows:
[0030] ;
[0031] In the formula, pr represents the priority, cs(n1, n2) represents the relevance, n1 and n2 represent the corresponding elements in any judgment vector, including task H, semantics S and behavior B, β is the weight coefficient, and its value is obtained by analyzing the association relationship model. The specific process is:
[0032] Historical statistics: Obtain historical correlation data and collect them, and mark them as set K = {K1, K2, ..., Ku}, where the correlation of each element in set K is sorted from low to high; then take values of set K every three digits, collect the data of the values, and mark them as set g;
[0033] Model calculation: Use the set g as the test data set and perform training and learning. Use the mean square error method to measure the linear relationship between the correlation and the weight, and automatically assign the weight.
[0034] Furthermore, interactive navigation includes task navigation, semantic navigation and behavioral navigation, and its content specifically includes: selecting different task menus, switching different scene contents and interacting with other elements in the corresponding virtual scene; wherein the other elements include virtual characters and virtual objects.
[0035] Furthermore, the step of building a preliminary calculation model includes: building a preliminary calculation model based on the initial rotation value and the guidance time, and combining the sight rotation value and the guidance time to generate a preset rotation value, based on the formula:
[0036] ;
[0037] Wherein, steer represents the preset rotation value, jd represents the sight rotation value, sj represents the guidance time, μ0 is the proportional coefficient, and the value range of μ0 is [0, 1].
[0038] Furthermore, the steps to adjust the preliminary calculation model are:
[0039] Proportion analysis: Feedback data represents the user's position feedback information for different virtual scenes, and the position feedback information includes small offset, medium offset and large offset. The proportion of users in small offset, medium offset and large offset is counted and marked as the first proportion, the second proportion and the third proportion respectively;
[0040] Comparative analysis: When users only have the first proportion, no processing is performed; when users have the second proportion or the third proportion, the preliminary calculation model is adjusted according to the difference between the corresponding proportion value and the corresponding standard threshold; wherein, the content of adjusting the preliminary calculation model is: adjusting the proportional coefficient μ0, and the formula based on the adjusted proportional coefficient is:
[0041] ;
[0042] Wherein, μ1 is the adjusted proportional coefficient, R is the error correction factor, and the value range of R is [0, 1], Zb represents the proportion value, and Bz represents the standard threshold.
[0043] (III) Beneficial effects
[0044] The present invention provides an automatic tracking system based on XR space, which has the following beneficial effects:
[0045] 1. By using a preset instruction library, the present invention can accurately map the identified eye posture to the specific analysis instructions in the XR space, obtain the gaze area and movement direction in the XR space, and can interact with the virtual reality device without the help of an external handle, which significantly improves the accuracy of the XR device and the user experience, making the user interaction in the XR space more natural, and providing technical support for more complex and delicate XR applications in the future; combined with the external environment, including light, temperature, humidity, flatness and network bandwidth, it plays a role in real-time correction in the process of solving the eye posture;
[0046] 2. The present invention performs rough and fine virtual scene division based on eye movement data, builds a first virtual scene and a second virtual scene, and can judge the virtual scene where the user is in real time, greatly improving the user's sense of immersion and participation; by establishing an association relationship model, it can perceive the user's attention and focus in real time, and accurately match this information with the virtual scene, and interact through eye movement input, and can set corresponding scenes for each user, so that it can be automatically adjusted according to the user's line of sight and focus area; to a certain extent, the adaptability of the virtual scene is improved;
[0047] 3. The present invention sets the initial rotation value and guidance time, and dynamically adjusts the preliminary calculation model according to user feedback, optimizes the guidance simulation, ensures that users have a realistic experience, and provides users in the XR space with a more personalized and customized experience. To a certain extent, it realizes accurate identification of the user's position in the virtual environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a module diagram of an automatic tracking system based on XR space according to an exemplary embodiment. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] Example
[0051] This embodiment provides an automatic tracking system based on XR space. Figure 1 is a schematic diagram of a module of an automatic tracking system based on XR space according to an exemplary embodiment. Figure 1 The system includes: a sight tracking module, a scene matching module, an interactive navigation module and a feedback adjustment module, and the sight tracking module, the scene matching module, the interactive navigation module and the feedback adjustment module are communicatively connected;
[0052] The following is a detailed description and explanation of each module:
[0053] Eye tracking module: automatically tracks the user's eyes;
[0054] Pre-build an analysis instruction library, map the eye posture to the corresponding analysis instruction in XR space based on the real-time tracked eye image, and identify the movement direction and gaze area of the eye;
[0055] The specific process is:
[0056] Image recognition: by quickly collecting eye images, based on the recognition of the frame difference images of the eyes, the frame difference images are differentially processed to obtain the first two frame difference images and the last two frame difference images, and the first two frame difference images and the last two frame difference images are ANDed to obtain the phase difference image;
[0057] The algorithm model for obtaining phase and difference images is:
[0058] ;
[0059] Where, d k (x, y, z) is the phase and difference image, b k-1,k (x, y, z) is the difference image of the first two frames, b k,k+1 (x, y, z) is the difference image of the last two frames, k is the number of eye image frames, and k is a positive integer greater than 2;
[0060] It should be noted that there are many ways to quickly collect eye images. For example, the user can use a glasses-type XR device with an adaptive built-in camera to collect the user's eye image. The device is not drawn in the figure. The collected eye image is usually a color image. By first processing the color image into a grayscale image and paying attention to the difference between the grayscale images, the differential image is processed based on the gray background image, and its movement or change area can be observed more intuitively.
[0061] State recognition: Using computer vision and machine learning algorithms, the eye movement trajectory points are captured based on the phase and difference images; the movement trajectory points are the normalized sequence of eye movement trajectory points;
[0062] It should be noted that the action trajectory point sequence can usually be described as a coordinate point in three-dimensional space. Assuming that the eyeball moves in a three-dimensional coordinate system, each action trajectory point can be represented as a position vector (xg, yg, zg); in practical applications, the action trajectory point is associated with time, and the position vector is represented as (xg_t, yg_t, zg_t), where t represents the time;
[0063] For example, the user's eye is at (0, 0, 0) at t=1, moves to (5, 5, 5) at t=5, and then moves to (10, 10, 10). The sequence of motion trajectory points is {(0, 0, 0, 1), (5, 5, 5, 5), (10, 10, 10, 10)}.
[0064] Motion mapping: Based on the motion trajectory points and pupil position, the eye posture is analyzed, matched with the predefined analysis instructions, and the eye movement direction and gaze area are analyzed and output;
[0065] The description of the pre-built analysis instruction library is as follows:
[0066] In order to ensure that the user's eye posture can be accurately identified and mapped to the correct analysis instructions, it is necessary to design and maintain a set of eye posture libraries and corresponding analysis instructions in advance. This eye posture library is the analysis instruction library, which contains common eye postures and their specific analysis instructions in the virtual environment.
[0067] The specific steps are:
[0068] Eye pose definition: A series of common eye images are identified and defined based on eye poses; for example, gaze and saccade; specifically:
[0069] The long axis radius and short axis radius of the eyeball are determined based on the motion trajectory points and the pupil position, and then combined with the environmental impact coefficient to analyze and calculate the eyeball posture;
[0070] Obtain the light, temperature, humidity, flatness and network bandwidth corresponding to the motion trajectory points, extract the environmental setting parameters in the database to obtain the set light interval, temperature interval, humidity interval, flatness interval and network bandwidth interval, standardize the light, temperature, humidity, flatness and network bandwidth and mark the processed results as: p1, p2, p3, p4 and p5 respectively;
[0071] By formula:
[0072] ;
[0073] In the formula, enc represents the environmental impact coefficient, , , , as well as are weight correction coefficients, , as well as It is a constant influencing parameter, usually taking a value of 0.5;
[0074] The algorithm model based on the eyeball posture is:
[0075] ;
[0076] Where, e is the eye posture solved based on the algorithm model, and the corresponding action is analyzed by matching based on the preset posture library; r max is the radius of the major axis, r min is the minor axis radius, g(·) is the morphological transformation function, then g(r max ) and g(r min ) is the major axis radius conversion value and the minor axis radius conversion value, which are used to analyze the shape of the eyeball. (d x , d y , d z ) is the center point of the phase and difference image, (r x , r y , r z ) is the pupil center, Φ is the eye rotation angle, I(u, v) is the pixel center of the phase difference image, is to find the integral value;
[0077] It should be noted that lighting affects the stability of the eye tracking system. Both strong light and low light may cause tracking errors, that is, under moderate lighting conditions, the eye position is most accurately acquired; temperature changes may affect hardware performance and sensor accuracy, thereby affecting the accurate acquisition of eye position; excessively high or low humidity may cause hardware failure or change facial reflection characteristics, affecting tracking accuracy, that is, under moderate lighting conditions, the eye position is most accurately acquired; a flat environment facilitates more stable head and eye position tracking, while an uneven environment may result in reduced tracking accuracy; low bandwidth or unstable network connections may cause delays or distortions in the transmission of eye position data, affecting the smoothness of the experience in XR space; by introducing an environmental impact coefficient, the eye position is corrected in real time when solving the eye position to ensure the accuracy of subsequent analysis and improve response time;
[0078] Command mapping: assigning specific analysis commands to each defined eye pose, which will control the user's gaze area and direction in XR space;
[0079] Training and calibration: The algorithm is trained using a large amount of sample data and calibrated through user testing to ensure the accuracy and responsiveness of motion direction and gaze area recognition;
[0080] For example: Gaze object: When the user's eye posture is aimed at a specific virtual object, the mapping instruction is "model interaction", at this time, the XR space triggers the corresponding interaction; Scan object: When the user's eye posture scans the XR space, the mapping instruction is "model loading", at this time, the XR space automatically loads related information or animation;
[0081] Through the above steps and the design of the eyeball posture library, efficient mapping of eyeball posture recognition and analysis instructions can be achieved, making the user's operation in the XR space more intuitive and natural; in this embodiment, the above steps also need to introduce the actual scene picture corresponding to the user's line of sight, based on the data information of all objects in the actual scene picture, so as to determine the position relationship between the line of sight area of the human eye in the XR space and the scene picture, laying the foundation for subsequent analysis;
[0082] Scene matching module: determines the virtual scene the user is in;
[0083] A first virtual scene is obtained by roughly dividing the scene based on the movement direction and the gaze area; a second virtual scene is obtained by finely dividing the scene based on the first virtual scene by introducing an attention mechanism, and an association relationship model is established based on the association relationship between the eye movement data and the corresponding virtual scene, and the association relationship includes a first association relationship and a second association relationship; wherein the eye movement data includes an eye movement trajectory, a gaze point, a scanning point, a movement direction and a gaze area;
[0084] The steps to determine the user's virtual environment are:
[0085] Rough identification: Divide the XR space into several different fields of view based on the gaze area and the movement direction. In each field of view, mark the user's gaze time and movement direction to generate high and low priority corresponding focus areas; load, unload or update the scene content for each focus area, and obtain the first virtual scene; the focus area includes the foreground area, the background area and the interaction area;
[0086] In XR space, in order to ensure efficient use of resources and improve user experience, different focus areas will be loaded, unloaded or updated. The content of which areas needs to be presented, updated or removed is determined based on the user's eye movement data and needs.
[0087] The specific steps are:
[0088] For high priority areas of concern:
[0089] Instant loading: Scene content needs to be loaded as early as possible to ensure that users can get a rich experience immediately when entering the area; for example, in VR games, game tasks and important UI information content will be loaded quickly and directly; Preloading: In order to avoid delays in scene loading, certain high-priority areas of interest can be preloaded, and the content of areas that users may enter can be loaded into memory in advance to ensure fast presentation when users need it; for example, in XR games, users may load the next task area in advance to reduce delays in loading;
[0090] Delayed unloading: Scene content is usually kept in memory until the interaction is completed or left; for example, in VR games, users switch task areas by switching gaze areas. When switching from one task area to another, the previous task area is unloaded; Automatic unloading: If memory resources are tight, the content of certain areas may be automatically unloaded when idle; for example, in XR games, the user has left a task area, and the scene content of that area is unloaded;
[0091] Real-time update: These areas need to be dynamically updated according to the user's operation to maintain the interactivity and immersion of the scene. They are usually the parts where users interact most frequently, such as virtual characters, mission objectives, interactive UI, dynamic scenes, etc. Priority cache update: During the interaction process, high-frequency updates are performed to keep the user smooth in the XR space;
[0092] For low priority areas of concern:
[0093] Lazy loading: loading is triggered only when the user's gaze area is close to these areas, which helps avoid wasting memory resources in areas that the user is not paying attention to and improves system performance; on-demand loading: content is loaded only when the user needs it. For example, if the user has left an area, the system can choose not to load the details or interactive content of that area temporarily to reduce memory usage.
[0094] Fast unloading: When the user's sight leaves these areas, the system can quickly unload them, such as auxiliary elements, decorative objects, and distant backgrounds in the scene; Memory optimization unloading: When the user quickly switches to the next area of interest, the scene content in these areas is completely unloaded to reduce unnecessary occupancy;
[0095] Periodic updates: updates are performed when the user approaches these focus areas, or even when the user does not focus on these areas at all; for example, the update frequency of some background elements and static objects can be reduced; on-demand updates: users update when system resources allow or only when the scene content changes; for example, when the user switches the scene perspective, some low-priority areas may reload or refresh their visual effects, but the update frequency is low and will not interfere with the main user interaction;
[0096] Fine recognition: Based on the first virtual scene, an attention mechanism is introduced to dynamically enhance the attention area of high priority, dynamically weaken the attention area of low priority, reconstruct the scene content, update the first virtual scene, and obtain the second virtual scene;
[0097] In XR space, in order to enhance the interactive experience and immersion in the virtual environment, the scene content will be dynamically enhanced or weakened. These operations usually involve dynamic management of scene content, and the visual effects in the virtual scene are adjusted in real time by continuously monitoring and analyzing the user's eye movement data, interaction, task priority and other information;
[0098] For dynamic enhancement:
[0099] Improved details: When the user approaches the area, the detail level of the corresponding scene content (such as texture resolution and model complexity) can be significantly improved; for example, in an XR game, when the user stands in front of an object, the details of the object will be enhanced in real time; Parallax enhancement: When the user interacts with an object in the XR space, the lighting, shadow and reflection effects of the object can be enhanced in real time to increase the presence and interactivity of the object; for example, in an XR space, there are many virtual characters. When the user interacts with a virtual character, the details and lighting effects of this object will be different from those of other characters; Physical simulation: When the user interacts with an object, the realism of the object will be increased; for example, the gravity, friction and elastic response of the object will be more realistic;
[0100] For dynamic weakening:
[0101] Fuzzification: Reduce visual saliency through fuzzification, so that users focus on high-priority areas; Interaction inhibition: Temporarily disable certain interactive functions and limit their application in the scene; for example, buttons and menus in certain low-priority areas no longer respond to user interaction actions; State change delay: Interaction response delay or slow feedback speed; for example, users' long-term operation will produce an effect;
[0102] By loading, unloading, updating, and dynamically managing scene content for each area of interest, it can effectively support interactive experiences in various virtual scenes;
[0103] The steps to establish the association between eye movement data and virtual scenes are:
[0104] Primary association: extracting the user's eye movement data based on the first virtual scene, creating a first time tag for the first virtual scene according to a time sequence, and establishing a first association relationship between the eye movement data and the first virtual scene;
[0105] Secondary association: based on the second virtual scene, extracting the eye movement data again, creating a second time tag for the second virtual scene according to the time sequence, and establishing a second association relationship between the eye movement data and the second virtual scene;
[0106] Model building: building a relationship model between the user and the virtual scene based on the first relationship, the second relationship, the first virtual scene, the second virtual scene, and the corresponding eye movement data;
[0107] The steps to guide users to interact in the XR space through the association model are:
[0108] Feature extraction: Use feature extraction technology to extract target features in the association model and construct feature vectors based on the target features, including tasks, semantics, behavior, and time;
[0109] It should be noted that feature extraction technology mainly relies on support vector machine SVM, natural language processing NLP, NEP technology, and LSTM technology; the meaning of the four-dimensional feature vector is explained as follows:
[0110] Task: indicates the task the user is performing, such as selecting an item or searching for an object; Semantics: indicates the specific meaning or situational information of the task, describing the context behind the task; Behavior: indicates the actual operation or behavior performed by the user, such as clicking, dragging, and rotating; Time: indicates the timestamp of the user's behavior in the XR space;
[0111] In-depth analysis: Preset the time period T. For any time period Ta, construct several different combinations of judgment vectors through feature vectors. The judgment vectors include: [task, semantics], [task, behavior] and [semantics, behavior]. Calculate the correlation between tasks, semantics and behaviors through cosine similarity, evaluate the importance of task-semantics, task-behavior and semantics-behavior, and set the corresponding priorities according to the importance to perform interactive navigation in sequence. Among them, task-semantics, task-behavior and semantics-behavior represent judgment combinations, which correspond to the judgment vectors [task, semantics], [task, behavior] and [semantics, behavior] one by one.
[0112] Among them, the correlation is calculated based on the cosine similarity, and the priority is calculated as follows:
[0113] ;
[0114] In the formula, pr represents the priority, cs(n1, n2) represents the relevance, n1 and n2 represent the corresponding elements in any judgment vector, including task H, semantics S and behavior B, β is the weight coefficient, and its value is obtained by analyzing the association relationship model. The specific process is:
[0115] Historical statistics: historical correlation data is obtained and collected, and marked as set K, K = {K1, K2, ..., Ku}, where the correlation of set K is sorted from low to high; the set K is valued every three digits, and the valued data is collected, and the set is g, then set g covers low, medium and high correlations;
[0116] Model calculation: The entire system is based on a deep network architecture. The set g is used as a test data set and input into the deep network for training and learning. During the training process, the mean square error method is used to measure the linear relationship between the correlation and the weight, automatically assigning a larger weight to a high correlation and a smaller weight to a low correlation;
[0117] For example, use NLP technology (such as Word2Vec, BERT) to convert tasks and semantics into embedded vectors, and mark them as H and S respectively, using the formula:
[0118] ;
[0119] In the formula, cs(H, S) represents the relevance, H·S represents the dot product between task and semantics, and Indicates the modulus length corresponding to the task and semantics; it should be noted that the value range of cs(H, S) is [0, 1], and its value is used to reflect whether the task execution meets the user's intention: 1 means that the correlation between the task and semantics is the highest, which means that the task execution and semantics are very consistent and the task can be completed effectively, while 0 means that the task and semantics do not match, and the user needs to be guided to make adjustments;
[0120] Similarly, calculate the correlation between the task and the behavior cs(H, B) and the correlation between the semantics and the behavior cs(S, B) in the time period Ta; by comparing the numerical values of cs(H, S), cs(H, B) and cs(S, B), set the larger the value, the greater the importance, and the smaller the value, the smaller the importance; for example: if the correlation between the task and the semantics is high, the task-semantics importance is high, and the task execution can be started directly; if the correlation between the task and the behavior is high, the task-behavior importance is high, and it can be automatically judged whether the task is in progress according to the user's behavior and give real-time feedback; if the correlation between the semantics and the behavior is high, the higher the semantic-behavior importance, it can be ensured that the execution of the task meets the user's expectations; then analyze the corresponding priorities to adjust the execution order between tasks, behaviors and semantics, and perform corresponding interactive navigation, the interactive navigation includes task navigation, semantic navigation and behavioral navigation, and its content specifically includes: selecting different task menus, switching different scene contents and interacting with other elements in the corresponding virtual scene; wherein, the other elements include virtual characters and virtual objects;
[0121] When the user's line of sight tracks a certain gaze area, interactive navigation is started and the user is guided to the corresponding location, including: Task navigation: guiding the user to complete the corresponding operation step by step according to the user's current behavior or task; for example, in XR games, providing prompts or directly guiding users to select items and use tools according to the sequence of game tasks; Semantic navigation: dynamically displaying related content or interactive options according to the user's current behavior or task, for example, in XR games, when the user is talking to the NPC, the NPC intelligently displays the corresponding documents and level resources according to the user's equipment; Behavioral navigation: automatically providing guidance according to the user's behavior and needs;
[0122] Feedback adjustment module: feedback and optional calibration of the user's position in the XR space;
[0123] The position of the virtual scene where the current user is located is identified, and a preliminary calculation model is built based on the initial rotation value and the guidance time under the condition that the user interacts according to the guidance to generate a real-time preset rotation value; then the user's feedback data is counted, and the preliminary calculation model is adjusted through analysis and processing, so as to perform recalibration of the user's position;
[0124] Position identification and initial rotation value setting: By analyzing the eyeball posture, the user's initial position (usually expressed in a coordinate system) and initial orientation (rotation value) in the virtual scene are obtained. These data provide the basis for the subsequent calibration process; Guidance time: Through the guidance time, the user's movement path within a specific time can be calculated. The path can be estimated through the user's movement trajectory or input data (such as handles, gait, sensor data, etc.) to form the user's movement trajectory;
[0125] The steps to build a preliminary calculation model are as follows: Based on the initial rotation value and the guidance time, build a preliminary calculation model, and combine the sight rotation value and the guidance time to generate a preset rotation value based on the formula:
[0126] ;
[0127] Wherein, steer represents the preset rotation value, jd represents the sight rotation value, sj represents the guidance time, μ0 is the proportional coefficient, and the value range of μ0 is [0, 1];
[0128] The steps to adjust the preliminary calculation model are:
[0129] Proportion analysis: Feedback data represents the user's position feedback information for different virtual scenes, and the position feedback information includes a small offset, a medium offset, and a large offset, which represents the difference or offset between the actual rotation value and the preset rotation value obtained by the user's actual operation. The settings of the small offset, the medium offset, and the large offset can be determined based on the interval. For example, the difference between the preset rotation value and the actual rotation value is calculated, and the difference range is set. When the difference is greater than 50%, it indicates a large offset. When the difference is between [5%, 50%], it indicates a medium offset. When the difference is less than 5%, it indicates a small offset. The specific setting method is set according to the actual situation and will not be described in detail here.
[0130] Count the proportion of users with smaller offset, medium offset, and larger offset, and mark them as the first proportion, the second proportion, and the third proportion respectively;
[0131] Comparative analysis: when the user only has the first proportion, no processing is performed; when the user has the second proportion or the third proportion, the preliminary calculation model is adjusted according to the difference between the corresponding proportion value and the corresponding standard threshold; wherein, the content of adjusting the preliminary calculation model is: adjusting the proportional coefficient μ0, and the formula based on the adjusted proportional coefficient is:
[0132] ;
[0133] Wherein, μ1 is the adjusted proportional coefficient, R is the error correction factor, and the value range of R is [0, 1], Zb represents the proportion value, and Bz represents the standard threshold value; wherein, the standard threshold value is set as follows: the corresponding historical proportion values are counted, and the corresponding average value and standard deviation are calculated, and the sum of the average value and 2 times the standard deviation is marked as the standard threshold value;
[0134] Thereby, the user position is recalibrated according to the adjusted preliminary calculation model.
[0135] In summary of the above technical solutions: the present invention includes a gaze tracking module, a scene matching module, an interactive navigation module and a feedback adjustment module; the technical points are: establishing a preset instruction library, accurately mapping the identified eye posture to the specific analysis instructions in the XR space, and obtaining the gaze area and movement direction in the XR space; performing rough and fine virtual scene division based on eye movement data, judging the user's virtual scene in real time, greatly enhancing the user's immersion and sense of participation, and then setting the initial rotation value and guidance time, and dynamically adjusting the preliminary calculation model according to user feedback, optimizing the guidance simulation, and ensuring that the user has a realistic experience.
[0136] In the application, the several formulas involved are all calculated by removing dimensions and taking their numerical values, and the formula is a formula obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The formula is set by technical personnel in this field based on actual conditions.
[0137] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. A person of ordinary skill in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0138] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected based on actual needs to achieve the purpose of the solution of this embodiment.
[0139] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. An automatic tracking system based on XR space, characterized in that: The system includes a gaze tracking module, a scene matching module, an interactive navigation module, and a feedback adjustment module; The eye tracking module identifies the eye position based on the real-time collected eye images, maps the eye position to the corresponding analysis instructions in the XR space based on the pre-built analysis instruction library, and identifies the movement direction and gaze area of the eye. A scene matching module obtains a first virtual scene based on a rough division of the motion direction and the gaze area; By introducing an attention mechanism, the first virtual scene is finely divided to obtain a second virtual scene, and an association relationship model is established based on the association relationship between the eye movement data and the corresponding virtual scene, and the association relationship includes a first association relationship and a second association relationship; wherein the eye movement data at least includes a movement direction and a gaze area; The interactive navigation module guides users to interact in the XR space through the association model; The feedback adjustment module identifies the position of the virtual scene where the current user is located, and builds a preliminary calculation model based on the initial rotation value and the guidance time under the condition that the user interacts according to the guidance, and generates a preset rotation value; then the feedback data of the user is counted, and the preliminary calculation model is adjusted through analysis and processing to perform recalibration of the user's position.
2. An automatic tracking system based on XR space according to claim 1, characterized in that: The steps to identify the direction of eye movement and the area of gaze include: Image recognition: by collecting eye images, identifying the frame difference images of the eyes, performing differential processing on the frame difference images, obtaining the first two frame difference images and the last two frame difference images, and performing AND operations to obtain the phase difference images; State recognition: Based on phase and difference images, computer vision and machine learning algorithms are used to capture the movement trajectory points of the eyeball. The movement trajectory points are the normalized sequence of the eyeball rotation trajectory points. Motion mapping: Based on the motion trajectory points and pupil position, the long axis radius and short axis radius are identified, and the eye posture is calculated in combination with the environmental impact coefficient. This is matched with the predefined analysis instructions to identify the eye movement direction and gaze area and output them. The steps for obtaining the environmental impact coefficient are as follows: collecting light, temperature, humidity, flatness and network bandwidth based on the motion trajectory points, and performing standardization processing according to the formula: ; Where enc represents the environmental impact coefficient, p1, p2, p3, p4 and p5 are the standardized light, temperature, humidity, flatness and network bandwidth. , , , as well as are weight correction coefficients, , as well as is a constant affecting parameter.
3. An automatic tracking system based on XR space according to claim 1, characterized in that: The steps of determining the virtual environment in which the user is located include: Rough recognition: Divide the XR space into several different fields of view based on the gaze area and movement direction. In each field of view, mark the user's gaze time and movement direction to generate high and low priority corresponding focus areas; load, unload or update the scene content in each focus area, and obtain the first virtual scene; the focus area includes the foreground area, the background area and the interactive area; Fine recognition: Based on the first virtual scene, an attention mechanism is introduced to dynamically enhance high-priority attention areas and dynamically weaken low-priority attention areas, and the scene content is reconstructed to update the first virtual scene to obtain the second virtual scene.
4. The automatic tracking system based on XR space according to claim 1, characterized in that: The steps of establishing the association relationship between the eye movement data and the virtual scene include: Primary association: extracting the user's eye movement data based on the first virtual scene, creating a first time tag for the first virtual scene according to a time sequence, and establishing a first association relationship between the eye movement data and the first virtual scene; Secondary association: based on the second virtual scene, extracting the eye movement data again, creating a second time tag for the second virtual scene according to the time sequence, and establishing a second association relationship between the eye movement data and the second virtual scene; Model building: Based on the first association relationship, the second association relationship, the first virtual scene, the second virtual scene and the corresponding eye movement data, a relationship model between the user and the virtual scene is built.
5. The automatic tracking system based on XR space according to claim 1, characterized in that: The steps to guide users to interact in the XR space through the association model include: Feature extraction: feature extraction technology is used to extract target features in the association model and construct a four-dimensional feature vector, including task, semantics, behavior, and time, which are marked as H, S, B, and time respectively; In-depth analysis: Preset the time period T. For any time period Ta, construct several different combinations of judgment vectors through feature vectors to calculate the correlation between tasks, semantics and behaviors, evaluate the importance of task-semantics, task-behavior and semantics-behavior, and set the corresponding priorities according to the importance to perform interactive navigation in sequence; among them, the correlation is calculated based on the cosine similarity, and the priority is calculated as follows: ; In the formula, pr represents the priority, cs(n1, n2) represents the relevance, n1 and n2 represent the corresponding elements in any judgment vector, including task H, semantics S and behavior B, β is the weight coefficient, and its value is obtained by analyzing the association relationship model. The specific process is: Historical statistics: Obtain historical correlation data and collect them, and mark them as set K = {K1, K2, ..., Ku}, where the correlation of each element in set K is sorted from low to high; then take values of set K every three digits, collect the data of the values, and mark them as set g; Model calculation: Use the set g as the test data set and perform training and learning. Use the mean square error method to measure the linear relationship between the correlation and the weight, and automatically assign the weight.
6. The automatic tracking system based on XR space according to claim 1, characterized in that: The steps of building a preliminary calculation model include: building a preliminary calculation model based on the initial rotation value and the guidance time, and combining the sight rotation value and the guidance time to generate a preset rotation value, based on the formula: ; Wherein, steer represents the preset rotation value, jd represents the sight rotation value, sj represents the guidance time, μ0 is the proportional coefficient, and the value range of μ0 is [0, 1].
7. The automatic tracking system based on XR space according to claim 1, characterized in that: Steps to adjust the preliminary computational model include: Proportion analysis: Feedback data represents the user's position feedback information for different virtual scenes, and the position feedback information includes small offset, medium offset and large offset. The proportion of users in small offset, medium offset and large offset is counted and marked as the first proportion, the second proportion and the third proportion respectively; Comparative analysis: When users only have the first proportion, no processing is performed; when users have the second proportion or the third proportion, the preliminary calculation model is adjusted according to the difference between the corresponding proportion value and the corresponding standard threshold; wherein, the content of adjusting the preliminary calculation model is: adjusting the proportional coefficient μ0, and the formula based on the adjusted proportional coefficient is: ; Wherein, μ1 is the adjusted proportional coefficient, R is the error correction factor, and the value range of R is [0, 1], Zb represents the proportion value, and Bz represents the standard threshold.
Citation Information
Patent Citations
Relative navigation method for eye movement interaction augmented reality
CN110285818A
Three -dimensional augmented reality operation navigation of self -adaptation developments based on real -time tracking and multi -information fusion
CN206649468U