Web interaction delay optimization

By performing semantic parsing and 3D spatial grid matching on multi-user voice communication data, preloading potential focus areas and generating guiding events, the problem of insufficient user interest point recognition in multi-user large-screen web interaction is solved, improving response speed and user experience.

CN120872208APending Publication Date: 2025-10-31ABA TIBETAN & QIANG AUTONOMOUS PREFECTURE ALL-ROUND TOURISM RESEARCH INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510970217.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In web interaction scenarios where multiple users share a large screen, existing technologies struggle to effectively capture the interests of multiple users, leading to biased focus area identification or idle resources, which impacts display efficiency and user experience.

Method used

By performing semantic analysis on multi-user voice communication data, a semantic vector model is constructed. Combined with the label information of the three-dimensional spatial grid region, potential focus areas are preloaded, and guided micro-interaction events are generated to induce the user's perspective to converge on the target area.

Benefits of technology

It achieves accurate prediction and response to user interests in a multi-user environment, improves system response speed and user immersion experience, adapts to complex interaction scenarios, and has self-optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872208A_ABST
    Figure CN120872208A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of image processing, and provides Web interaction delay optimization, which comprises the following steps of: performing space grid division on a target area in a Web end three-dimensional digital space, and constructing a space index structure; collecting voice communication data among the plurality of users, and constructing a semantic vector model of a target currently concerned by the plurality of users; target areas possibly interested by the multiple users subsequently are determined to serve as a potential focus area set; performing content resource preloading on the potential focus area set, generating a guided micro-interaction event on the potential focus area, and inducing view angles or interaction behaviors of the plurality of users to gather towards a preloading area; and displaying the preloaded content on the large screen based on the interactive operation of the plurality of users on the focus area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing, and specifically relates to a method for optimizing web interaction latency. Background Technology

[0002] With the development of 3D modeling, WebGL graphics rendering, and high-performance browser engines, web-based 3D interactive display technology has been widely applied in fields such as smart tourism, digital museums, virtual tours, and ecological science popularization. By deploying large-size, high-resolution display terminals combined with 3D digital scene models, users can immerse themselves in browsing rich text, images, videos, and dynamic simulation content, enhancing spatial awareness and cultural dissemination. Especially in cultural theme exhibition halls, web-based 3D displays have become the mainstream interactive medium.

[0003] To optimize the user browsing experience and reduce latency and stuttering caused by content loading, existing technologies generally employ preloading mechanisms to load content to be accessed in advance. Common methods mainly fall into two categories: first, eye-tracking technology, which uses infrared cameras or posture recognition modules to monitor the user's eye movements or head orientation to predict the content the user is about to view; second, user behavior modeling, which statistically models the user's click records, swipe paths, or areas of inactivity within a scene to determine their interest trends and load corresponding resource files accordingly. These technologies have achieved some success in personal devices such as AR / VR glasses, mobile phones, or large-screen devices operated by a single user.

[0004] However, existing methods have significant shortcomings in interactive scenarios involving multiple users sharing a large screen. In such environments, multiple users can simultaneously participate in voice discussions, view displayed content, or express interest, but the large screen system only responds to a single main control entry point, such as the operator's actions or explanation trajectory. When the system preloads content based solely on a single-user prediction mechanism, it struggles to effectively capture the interests of most users, leading to biased or missing focus areas. Furthermore, without a guidance mechanism, even if the system has completed resource preloading, users may fail to enter the relevant area due to a lack of effective visual guidance, resulting in idle preloaded resources, missed hot content, and a fragmented user experience, ultimately reducing display efficiency and content engagement. Summary of the Invention

[0005] To address the problems in existing technologies, this invention provides a web interaction latency optimization method for displaying 3D digital scenes on large-screen devices, where the large screen simultaneously displays content to multiple users. The method includes the following steps:

[0006] The target area in the three-dimensional digital space of the Web terminal is divided into spatial grids, and a spatial index structure is constructed. The spatial index structure is used to describe the spatial position relationship and visual attributes of each grid area in the three-dimensional scene.

[0007] Collect voice communication data between the multiple users, perform semantic parsing on the voice data to extract keywords, intent expressions and contextual semantic information, and construct a semantic vector model of the current focus of the multiple users;

[0008] The semantic vector model is semantically matched with the label information of each grid region in the three-dimensional digital space to determine the target regions that the multiple users may be interested in later, as a set of potential focus regions.

[0009] Content resources are preloaded onto the set of potential focus areas, and guided micro-interaction events are generated on the potential focus areas to induce the perspectives or interaction behaviors of the multiple users to converge on the preloaded areas.

[0010] Based on the interactive operations of the multiple users on the focus area, the preloaded content is displayed on the large screen.

[0011] Furthermore, the spatial grid is divided into regular cubes or deformed tetrahedral structures, and a spatial index structure is constructed that includes node positions, adjacent topological relationships, visible range, and occlusion attributes.

[0012] Furthermore, the voice communication data is collected by an omnidirectional microphone array installed at the top of the display area or on the desktop, and background noise and non-voice signals are filtered out by a noise suppression algorithm.

[0013] Furthermore, the process of constructing the semantic vector model includes inputting the text transcribed from speech recognition into a natural language processing model based on a deep neural network, and outputting a semantic embedding vector of fixed dimensions. The natural language processing model is at least one of Word2Vec, GloVe, or BERT.

[0014] Furthermore, the tag information is a set of semantic tags preset for each grid area when constructing a 3D scene. The tag set includes scenic spot name, historical background, cultural attributes, geographical features, or ecological theme text information.

[0015] Furthermore, the semantic matching uses a cosine similarity algorithm to calculate the similarity between the semantic vector and the label vector, and sets a matching threshold or selects the top N regions in terms of similarity to form a set of potential focus regions.

[0016] Furthermore, the preloading of the content resources includes 3D model files, texture layers, multimedia materials, and interactive scripts, and adopts an asynchronous loading mechanism and a layered loading strategy, prioritizing the loading of low-resolution resources.

[0017] Furthermore, the guided micro-interaction events include generating dynamic icons, glowing boundaries, indicating paths, or voice prompts in the focus area to guide the user's viewpoint toward the target area or trigger further interaction.

[0018] Furthermore, the user's interactive operations include voice commands, mouse clicks, gesture pointing, long-term viewing pauses, or laser pointer trajectory input. These operations are identified in real time by the event listening module and mapped to the spatial index structure.

[0019] Furthermore, after the pre-loaded content is displayed on the large screen, the system also records and analyzes the user's actual response behavior to update the regional heat map, optimize the resource scheduling order, and adjust the guidance event generation strategy.

[0020] This invention collects voice communication data between multiple users, extracts semantic information to construct a group attention intent vector model, and combines it with the semantic label of a three-dimensional spatial grid region for matching to determine potential focus areas. This solves the problem that traditional methods are difficult to determine user attention targets in multi-user collaborative browsing scenarios, and achieves accurate content prediction for multiple users.

[0021] This invention introduces a preloading mechanism and micro-interaction guidance strategy based on potential focus areas. After the resources are loaded, the system actively attracts the user's attention to the target area through visual cues, dynamic markings, and voice reminders. This not only improves the system's response speed but also significantly enhances the user's immersive experience and interaction efficiency.

[0022] Furthermore, this invention constructs a closed-loop process from semantic parsing, interest inference, preloading scheduling, interactive response to display feedback, which can adapt to complex interactive scenarios and achieve self-optimization by recording users' actual operation behavior. It has good adaptability and scalability and is suitable for various large-scale display system scenarios such as digital exhibition halls, virtual tourism, and remote collaboration. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram illustrating the principle of Web interaction latency optimization in this invention. Detailed Implementation

[0025] The invention will now be described in preferred form with reference to the accompanying drawings and specific embodiments.

[0026] This embodiment solves the above problems through the following steps:

[0027] In one embodiment, a Web interaction latency optimization method addresses the issue of delayed user interaction response when displaying 3D digital scenes in a Web environment. This method involves constructing a spatial index structure, integrating multi-person voice semantic analysis to determine the user's potential focus area, preloading relevant resources, and using micro-interaction guidance techniques to induce the user's operating perspective to migrate to a low-latency area. This significantly reduces the perceived latency during the interaction process and enhances the user's response experience and immersion.

[0028] The Web platform is used to display 3D digital scenes on large-screen devices. These large screens are high-definition display terminals permanently installed in exhibition spaces, government service halls, cultural exhibition centers, or smart tourism experience centers. They are capable of simultaneously outputting images, audio, and interactive prompts to multiple users, making them suitable for scenarios where multiple users can browse, interact, and receive guided explanations in public areas. The 3D digital scenes are constructed based on the needs of Aba Prefecture's digital urban and rural construction, encompassing virtual reconstructions of natural landscapes, cultural relics, artifacts, village distribution, ecotourism routes, and digital rural infrastructure, and are presented through a Web-based 3D visualization engine.

[0029] The three-dimensional digital scene is an immersive interactive virtual environment built on the web, which runs in a browser environment using WebGL, Web3D or other compatible visualization technologies, and presents three-dimensional content including buildings, blocks, cultural relics, equipment or abstract models.

[0030] In the aforementioned application scenarios, multiple users can simultaneously operate the 3D scene via voice, gestures, remote control, mobile terminals, or other interactive methods, such as rotating the viewpoint, selecting hotspots, retrieving information, and switching scenes. However, due to limitations in network transmission bandwidth, browser processing performance, and 3D rendering overhead on the web interface, issues such as response delays, screen stuttering, or untimely content loading are prone to occur during interactive operations, especially when there are a large number of concurrent users and the target scene is highly complex.

[0031] For example, in the Aba Prefecture ecotourism-themed exhibition hall, the web system is used to display 3D digital landscape scenes of the Jiuzhaigou area on a large screen. The display integrates models of multiple tourist attractions with real-world geographical locations, including but not limited to representative areas such as Nuorilang Waterfall, Shuzheng Lakes, Wuhua Lake, and Wucai Pond. Visitors can give switching commands via voice interaction, such as "view Wucai Pond." The system needs to interpret the user's intent based on the voice content and switch the perspective and load local layers in the current 3D scene. However, because the 3D model of the Jiuzhaigou panoramic area contains a large amount of high-precision terrain data, texture information, and detail layers, the web client often encounters problems such as slow screen switching, blurry image loading, or delayed voice command response when interpreting commands and completing area switching due to the large amount of data and lack of pre-loading of resources. This seriously affects the user's immersion, continuity, and interaction satisfaction with the digital cultural tourism content.

[0032] Therefore, to address the latency issues caused by the inconsistency between resource loading and interactive feedback in the aforementioned web-based operating environment, there is an urgent need for an optimization mechanism that can dynamically sense user interaction intentions, predict target areas, and preprocess resources in advance, in order to improve the response efficiency and experience quality of web interactive systems in large-screen, multi-user environments.

[0033] To solve the above problems, one embodiment of the present invention is implemented through the following steps, the implementation principle of which is as follows: Figure 1 As shown:

[0034] Step S10: Divide the target area in the Web terminal 3D digital space into spatial grids and construct a spatial index structure. The spatial index structure is used to describe the spatial position relationship and visual attributes of each grid area in the 3D scene.

[0035] In web-based 3D digital display systems, 3D scenes typically contain numerous geographic elements, architectural models, texture layers, and user interaction hotspots. To improve the positioning accuracy of user interactions and the system's response speed to areas of user interest, it is necessary to logically divide and structurally manage the entire 3D space. By dividing the 3D space into several regular or irregular spatial grids and constructing corresponding spatial index structures, rapid positioning, semantic annotation, loading control, and viewpoint guidance for any user-operated area can be achieved. This effectively reduces content retrieval and response latency, improves resource scheduling efficiency, and enhances the system's real-time performance.

[0036] In this invention, spatial grid partitioning refers to discretizing three-dimensional space according to predetermined geometric rules (such as cubes, hexahedrons, or tetrahedrons) to form multiple sub-regions with clearly defined boundaries, called grid cells. Each grid cell has a unique coordinate index and spatial range in three-dimensional space. The spatial index structure refers to the logical structure used to store, manage, and query the aforementioned grid cells. It can be a tree structure, a hash structure, or a hierarchical structure built based on spatial partitioning rules, used to quickly determine the spatial range or region attribute to which a user operation belongs.

[0037] In one optional implementation, the step of dividing the target region in the three-dimensional digital space on the Web platform into a spatial grid and constructing a spatial index structure specifically includes the following sub-steps:

[0038] Step S101: Determine the boundary range of the three-dimensional digital space. Obtain the three-dimensional coordinate boundary parameters based on the actual size of the display scene, including the maximum and minimum values ​​of the X-axis, Y-axis and Z-axis, and form a spatial bounding box to define the divisible effective display area.

[0039] Step S102: Set the mesh granularity for spatial division. Determine the size of the division unit based on the complexity of the 3D scene content, the user's interaction accuracy requirements, and system resource limitations. For example, a smaller mesh granularity can be used in areas displaying dense models, while a larger mesh can be used in open scene areas to achieve non-uniform division.

[0040] Step S103: According to the set granularity, perform regular meshing on the entire three-dimensional space to generate multiple mesh units. Each unit is a hexahedral voxel, and its three-dimensional coordinate range and unique index number in the global space are recorded.

[0041] Step S104: Extract the 3D model information, visual attributes and semantic identification information contained in each grid cell. The visual attributes include the rendering complexity of the area, the number of textures, whether it is a user-visible area, etc. The semantic identification information includes the scenic spot name, function type or cultural meaning of the area.

[0042] Step S105: Construct a spatial index structure using an octree to efficiently organize and retrieve grid cells. Each node records the spatial region it contains and its subordinate child nodes. For dynamically loaded scenarios, status flags can be added to indicate whether the region has been loaded, whether it is a hotspot region, etc.

[0043] Step S106: Bind the spatial index structure to the subsequent user behavior analysis module via an interface, so that when a user interaction event occurs, the corresponding grid cell can be quickly located and its visual attributes and semantic tags can be extracted for subsequent resource scheduling and preloading mechanisms to call.

[0044] The above method manages the three-dimensional spatial content in a structured way, enabling the system to identify potentially involved spatial areas and schedule resources in advance before user operation, thus avoiding the computational redundancy and delay risks caused by the system processing the entire three-dimensional scene indiscriminately.

[0045] Taking the Jiuzhaigou 3D digital display in the Aba Prefecture Ecotourism Theme Exhibition Hall as an example, the system first establishes a spatial bounding box for the entire Jiuzhaigou 3D model, defining the scene boundary range as X-axis -200 meters to +200 meters, Y-axis -150 meters to +150 meters, and Z-axis 0 meters to +100 meters. Then, based on the density of attractions, the entire area is divided into a cubic grid with each side length of 10 meters, and the attraction models, texture information, and geographic tags within each grid are extracted and labeled. Finally, the system establishes a spatial index structure based on an octree, so that when a visitor issues the voice command "view the Five-Colored Pool," the system can quickly locate the grid area where the Five-Colored Pool is located, complete resource retrieval and content preloading, ensure visual continuity and response consistency during the interaction process, and thus improve the interactive performance and experience quality in a multi-user shared display environment.

[0046] Step S20: Collect voice communication data between the multiple users, perform semantic parsing on the voice data to extract keywords, intent expressions and contextual semantic information, and construct a semantic vector model of the current focus of the multiple users.

[0047] In multi-user, 3D web interactive environments, voice communication is the primary way for users to express their operational intentions, convey content of interest, and coordinate instructions. Since large-screen display systems are used by multiple users simultaneously, traditional single-user click-based interaction mechanisms are insufficient to meet the demands for efficiency and naturalness in concurrent interactions. Therefore, to improve the system's responsiveness to user needs, it is necessary to collect and analyze voice communication data between users in real time, extract users' current intent information through semantic understanding technology, and thus construct a multi-user interest and attention model to support subsequent content prediction, preloading, and interactive guidance.

[0048] In this invention, voice communication data refers to the audio stream formed by users' conversations, commands, or questions in natural language within the display space, which can be collected through microphone arrays or directional sound pickup devices. Semantic parsing refers to the linguistic and semantic analysis of the text information obtained after voice signal recognition, extracting keywords, action intentions, target objects, and contextual logical relationships. The semantic vector model encodes semantic information into a vector form that can be understood and processed by computers, using multi-dimensional feature vectors to express the semantic features of the target area or topic content currently of interest to the user, for matching and positioning in three-dimensional space.

[0049] In one optional implementation, the step of collecting voice communication data between the multiple users and performing semantic parsing to construct a semantic vector model of the target currently of interest to the multiple users specifically includes the following sub-steps:

[0050] Step S201: Collect voice communication data of multiple users in the display area. The voice data is picked up in real time by an omnidirectional microphone array installed on the top, wall or table of the display space. Each microphone array corresponds to a pickup channel and can optionally be configured with an array beamformer to enhance the voice signal in a specified direction. At the same time, the system performs frame-level segmentation processing on the collected raw audio stream and calls noise suppression algorithms (such as spectral subtraction, Wiener filter or deep learning-driven DNN voice enhancement model) to filter out background ambient noise, echo signals and non-voice events (such as footsteps, door opening sounds, etc.) to generate clear human voice audio segments.

[0051] Step S202 involves performing speech recognition processing on the human voice audio segment using an automatic speech recognition engine based on a deep neural network structure to convert the audio signal into analyzable text content. Optional implementation schemes include: using the DeepSpeech model, which constructs a mapping from speech to character sequences based on the CTC loss function; or using the wav2vec2.0 model, which pre-trains the audio segment to text encoding through self-supervised learning. The recognition result includes the complete sentence text, speech start and end timestamps, and the user identifier corresponding to the audio channel (e.g., MIC_01, MIC_02), and is stored as structured data for subsequent processing.

[0052] Step S203 involves performing linguistic structure analysis on the identified text content. Natural language processing tools (such as spaCy, Stanza, or HanLP) are used to perform word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing to extract keywords (such as "Five-Colored Pool"), verb predicates (such as "display," "switch"), modal or interrogative structures (such as "can / cannot," "please ask"), and related noun phrases and target object descriptions (such as "autumn scenery," "Five-Colored Pool in the Jiuzhaigou area"). A grammatical dependency graph is constructed to clarify the syntactic connections between the subject, predicate, and object of each statement, thereby identifying the user's core operational intent.

[0053] Step S204 involves performing conversation-level context fusion on the voice texts of multiple users. A sliding time window mechanism is used to merge and align multiple statements within the same time period, identifying semantically common keyword combinations or high-frequency target phrases. For example, when multiple users successively send "Show Five-Colored Pool," "I want to see the section about Five-Colored Pool," and "Is there a Five-Colored Pool in autumn?", the system uses keyword merging and intent clustering algorithms to determine that "Five-Colored Pool" is the most focused target content in the current conversation, thus identifying it as the group semantic focus.

[0054] Step S205: Based on the system's preset 3D digital scene labeling system, the aforementioned keywords or target phrases are matched with the labeled semantic tags in the scene. A word embedding model is used to map the natural language target to a vector space, forming a structured semantic vector model. Optional implementation schemes include: using the Word2Vec model to convert "Five-Colored Pool," "autumn scenery," and "right side of Jiuzhaigou" into 300-dimensional word vectors, and weighting similar word vectors to form a weighted average to form the overall semantic vector V_user; when using the BERT model, the entire sentence "showing the Five-Colored Pool in autumn" can be encoded, outputting a 768-dimensional context-related semantic vector V_bert. The semantic vector model is represented as follows:

[0055] V_focus=w1*V_1+w2*V_2+...+wn*V_n,

[0056] Here, V_i represents the semantic vector of each keyword or phrase, and w_i represents its weight or frequency weight in the current voice communication. This semantic vector V_focus serves as an aggregated representation of the user's current group attention intent in the semantic space.

[0057] For example, when a user's voice contains the following text: "Let's see what the Five-Colored Pond looks like in autumn," "Is it on the right?" "Is there an introduction?", the system extracts the semantic vector of each sentence using the BERT model and merges them into the final vector:

[0058] V_focus ≈ BERT("Five-Colored Pool") + 0.7 * BERT("Autumn") + 0.3 * BERT("Right Side")

[0059] This information is used to perform similarity calculations (such as cosine similarity) with the semantic label vectors of each grid region in 3D space, thereby determining the regions that the user group may be interested in.

[0060] Step S206 involves inputting the constructed semantic vector model into the aforementioned spatial index structure. By comparing the semantic vector with the semantic label vectors defined in the grid cells, a similarity score is calculated, and several of the most relevant potential focus areas are determined based on the score ranking. The system triggers resource preloading, scene activation, and visual guidance mechanisms for the areas with the highest matching degree, enabling the rendering preparation of relevant area content to be completed before the user issues an explicit operation command, thus improving response speed and interaction consistency.

[0061] Through the above steps, the system can accurately identify and predict the content that users are interested in based on semantic modeling during multi-user parallel conversations.

[0062] For example, in the Aba Prefecture Ecotourism Theme Exhibition Hall, when multiple visitors are talking about the three-dimensional scene of Jiuzhaigou, keywords such as "Five-Colored Pool," "autumn," "red leaves," and "right side" appear repeatedly in their voices. The system constructs these keywords into a unified semantic vector model according to the above processing flow and matches them to the grid area corresponding to "Five-Colored Pool" in three-dimensional space, thus helping to complete the content preparation before the user issues a clear switching instruction.

[0063] Step S30: Semantically match the semantic vector model with the label information of each grid region in the three-dimensional digital space to determine the target regions that the multiple users may be interested in later, as a set of potential focus regions.

[0064] In 3D web display systems, user interactions often involve expressing their interests in natural language, while the 3D space is composed of several grid regions with labeled information. To enable the system to quickly respond to and load corresponding content based on the user's voice intent, it is necessary to semantically match the previously constructed semantic vector model with the label information corresponding to each grid region in the 3D digital space. By comparing the semantic vector with the grid label vectors, a set of spatial regions closest to the user's current interest can be identified as a potential focus area set, thus providing a basis for subsequent content preloading and interactive guidance.

[0065] The preceding steps yielded a semantic vector model, which is a vector form encoded by a natural language processing model, representing the user's intent expressed through their speech semantic content. This model reflects the user's center of interest in the semantic space. Grid label information refers to the pre-labeled descriptive information for each grid region in three-dimensional space, including the region's geographical location, cultural attributes, and functional type. This information is recorded in natural language and converted into a label semantic vector using word vector technology, used for similarity comparison with the user's semantic vector.

[0066] In one optional implementation, the step of semantically matching the semantic vector model with the label information of each grid region in the three-dimensional digital space and determining the set of potential focal regions specifically includes the following sub-steps:

[0067] Step S301: Extract label information for all grid regions in the 3D digital space. This label information is manually configured or automatically extracted during the 3D scene modeling stage and bound to each grid cell to describe the entity or semantic features represented by that grid. The label information is represented in natural language and includes, but is not limited to: the corresponding geographical identifier (e.g., "Huanglong Temple," "Hongyuan Grassland"), seasonal attributes or ecological characteristics (e.g., "Autumn Landscape," "Spring Flower Sea"), cultural and historical elements (e.g., "Tibetan and Qiang Folk Customs," "Tea Horse Road Relics"), functional uses (e.g., "Visitor Service Center," "Main Entrance Passage"), and other textual descriptions that can represent the scene's semantics. This label information is stored in a structured metadata table and corresponds one-to-one with the grid numbers in the spatial index structure, forming a label mapping matrix.

[0068] Step S302 involves text encoding the tag information, converting each tag text into a fixed-length semantic vector using a natural language processing model. To achieve this semantically numerical representation, possible implementation schemes include:

[0069] The Word2Vec model maps words to dense vectors, making it suitable for processing word-level labels (such as "lake" and "valley").

[0070] Using the GloVe model, word vectors are constructed based on global corpus statistical features, which is suitable for scenic spot categories with broad semantic distribution;

[0071] The FastText model has a strong ability to identify word roots and affixes, and is suitable for tags containing place names, compound words, etc.

[0072] Using a BERT pre-trained model to encode the entire sentence while preserving contextual dependencies, it is suitable for processing complete phrase labels (such as "autumn scenery of the Five-Colored Pond").

[0073] The encoding process includes word segmentation, stop word removal, word vector generation, and vector normalization. The final output tag vector is denoted as V_tag(i), where i represents the i-th grid region, corresponding to dimension d (e.g., d = 300 or 768). For grid regions with multiple tags, a weighted average method can be used to generate a comprehensive tag vector.

[0074] V_tag(i)=(w1*V1+w2*V2+...+wn*Vn) / (w1+w2+...+wn)

[0075] Where V1Vn represents the word vector of each sub-tag, and w1wn is the weight of each sub-tag, which is usually set according to the importance of the tag or the confidence of the source.

[0076] Step S303 involves calculating the semantic similarity between the previously constructed user semantic vector model V_focus and the tag vector V_tag(i) of each grid region. This similarity is used to measure the degree of matching between the user's currently focused content and the semantics of each grid. The similarity calculation uses the standard cosine similarity formula:

[0077] sim(i)=(V_focus·V_tag(i)) / (||V_focus||×||V_tag(i)||)

[0078] in:

[0079] sim(i) is the similarity score between the user's semantic vector and the semantic vector of the i-th grid.

[0080] "·" represents the dot product of vectors;

[0081] ||V_focus|| represents the Euclidean norm (L2 norm) of the user semantic vector;

[0082] ||V_tag(i)|| represents the Euclidean norm of the tag semantic vector.

[0083] After the calculation is completed, the system generates a similarity distribution list equal to the number of grid cells, denoted as S={sim(1),sim(2),...,sim(N)}, where N is the total number of grid cells.

[0084] Step S304: Based on the similarity score in the semantic matching results, potential focus regions are filtered. Optional filtering strategies include:

[0085] Set a fixed threshold θ (e.g., θ = 0.80), and select all grid regions where sim(i) ≥ θ to form a potential focus set R_focus;

[0086] The top N highest similarity scores are set as the selection criteria, that is, the grid numbers corresponding to Top_N{sim(i)} are selected to enter R_focus.

[0087] The selected regions can be further aggregated. If several highly similar regions are adjacent or close to each other in space, they can be merged into a logical focus block for subsequent unified loading and guidance.

[0088] Step S305: Add state marker information to each grid region in R_focus, including but not limited to the following:

[0089] Loading status flag: Indicates whether the 3D models, textures, and script resources for this area have been loaded;

[0090] First-time recognition marker: Indicates whether this area is the first time it has been recognized as the focus during this voice interaction;

[0091] Viewpoint association marker: Indicates whether the area is within the current user's main view's visible range;

[0092] Historical access frequency: Records the frequency with which this area has been accessed or mentioned in past interactions, serving as a reference for the strength of smart guidance.

[0093] The aforementioned status information is stored in a structured form in the focus area management module and serves as the basis for determining the priority of guided micro-interaction generation and resource preloading in subsequent steps.

[0094] Through the above refined processing flow, the system can accurately locate multiple potential areas of interest in a complex 3D display space based on the user's natural language intent, and complete scene preparation in the background, providing a strong response guarantee for subsequent user active interaction, effectively solving the problems of response latency and sudden changes in resource loading in 3D Web systems.

[0095] For example, in the panoramic display scene of Jiuzhaigou in the Aba Prefecture Ecotourism Theme Exhibition Hall, visitors frequently mentioned keywords such as "autumn," "Five-Colored Pool," "red leaves," and "scenery" through voice. The system constructed a user semantic vector V_focus, and after cosine matching with the label vectors of each grid area, selected grid numbers G_172, G_173, and G_174 with a similarity greater than 0.85, and merged them into a focus area block. At the same time, it was marked as "not loaded, first appearance, visible from the main viewpoint." The system then performed texture loading and guided generation in this area to prepare for the upcoming user operation.

[0096] Step S40: Preload content resources into the set of potential focus areas, generate guided micro-interaction events in the potential focus areas, and induce the perspectives or interaction behaviors of the multiple users to converge into the preloaded areas.

[0097] In a large-screen 3D digital display environment shared by multiple users, the content is presented to multiple users simultaneously, and voice interaction is usually participated in by multiple users. However, the actual control perspective or execution of operation commands is often concentrated on a single main operator, such as a presenter, lecturer, or guide. In such scenarios, the main operator may not fully pay attention to the voice discussions among other visitors, especially in unstructured, natural language-heavy contexts, which can easily lead to information asymmetry and result in some users' needs not being addressed in a timely manner.

[0098] Therefore, when the system identifies certain users who frequently mention specific scenic spots or areas (such as "Five-Colored Pool", "Autumn Scenery", "Red Leaves") through semantic analysis, even if the main operator has not yet switched the view to that area, the system should still actively generate guided micro-interaction events to make that area visually change significantly on the large screen, attracting the main operator's attention in a non-intrusive way, thereby realizing the transformation from passive response to active guidance.

[0099] The core of this mechanism is to use multi-user semantic consensus to identify potentially important areas, and then convey the importance of these areas to the actual controllers through visually induced feedback, triggering their interactive behavior. This allows the main operator to subconsciously receive the signal that "this area is of concern to others," thereby completing an autonomous perspective switching operation and achieving a coordinated response of collective interests.

[0100] Step S401: Obtain the resource list information for each grid region in the potential focus region set R_focus. The resource list is generated by the system or configured by the developers during the 3D scene deployment phase, and the list information includes, but is not limited to, the following resource items:

[0101] 3D model files: used to construct the spatial geometry of the mesh region, usually in formats such as .obj, .fbx, .glb, .gltf, etc., containing coordinate vertices, normals, UV information, etc.

[0102] Texture layer files: used to map and display terrain, vegetation, water surfaces or building appearances. File formats include .jpg, .png, .dds, etc., and support multi-layer texture overlay;

[0103] Scripts and interaction logic files: These are used to define the behavioral logic of user interaction with the scene, such as click response, hotspot trigger, video playback call, etc. They are often in .js, .ts or blueprint script format;

[0104] Multimedia materials include high-definition videos (such as .h264, .mp4), surround sound effects (.ogg, .mp3), narration audio files, background music, etc.

[0105] Annotations and information display files: including HTML-formatted text descriptions, graphic introductions, related link data, etc.;

[0106] Additional information such as data loading path, caching strategy parameters, and resource version identifier.

[0107] The resource list is bound to the spatial index node of each grid, forming a one-to-one mapping relationship between the index and the resource list, which can be quickly retrieved during subsequent asynchronous scheduling.

[0108] Step S402: Determine the current resource status of each grid region in R_focus and determine the preloading target and priority. The resource status is maintained by the system during runtime and includes, but is not limited to:

[0109] Loading status flags (not loaded, partially loaded, ready);

[0110] Cache validity flag (whether it exists in the local cache);

[0111] Last access time and current activity level (used to determine the cold / hot status of resources).

[0112] The system prioritizes users based on the three-dimensional Euclidean distance D(i) between the user's current position P_user and the center P_i of each focal region:

[0113] D(i)=sqrt[(x_user-x_i)^2+(y_user-y_i)^2+(z_user-z_i)^2]

[0114] Where (x_user, y_user, z_user) are the coordinates of the user's current virtual viewpoint, and (x_i, y_i, z_i) are the coordinates of the center point of the i-th grid region.

[0115] Regions with higher priority will be added to the resource scheduling queue first. The scheduling strategy uses a multi-threaded asynchronous loading mechanism, with the main thread maintaining page responsiveness and worker threads performing resource fetching operations. For web environments, Three.jsloader or Babylon.jsResourceManager can be used, while UE4's PixelStreaming resource loader can be used for high-performance engines.

[0116] Step S403 involves employing compression and block decompression strategies during resource loading to optimize loading efficiency and reduce bandwidth consumption. Specific implementation includes:

[0117] The 3D model uses a simplified LOD loading method. That is, a low-precision model such as LOD1 or LOD2 is loaded to build the general outline, and a high-precision LOD0 model is loaded when the user approaches or the view is locked.

[0118] Texture layers are loaded on demand using a tile-based mechanism. For example, a quadtree-based tile hierarchy structure can be used, loading low-resolution layers first, followed by progressively loading high-resolution tiles.

[0119] The video is transmitted using streaming media protocols such as HLS (HTTP Live Streaming) or MPEG-DASH, which divide the video into several segments. Each segment is prefetched before it is needed for playback, avoiding loading the entire file at once.

[0120] Use gzip or Brotli compression algorithms for transmitting scripts and multimedia materials to reduce data packet size.

[0121] After the resources are loaded, the system updates the region status to "ready" and records the corresponding resource identifier in the local cache index.

[0122] Step S404: Determine whether a guided micro-interaction event should be triggered based on the relationship between the focus area and the user's current viewpoint. The system first determines whether the target area is within the projectable range of the user's main viewpoint based on the current view frustum. If it is within the range, the system calculates its screen projection coordinates and determines the saliency score.

[0123] The significance score is calculated based on the following factors:

[0124] Screen position (the closer to the center, the higher the weight);

[0125] Area size;

[0126] The degree of dynamic background interference in the current scene;

[0127] Historical click / visit frequency.

[0128] If the salience score exceeds a preset threshold, a micro-interaction event is triggered. Optional event types include:

[0129] Visual cues: such as a breathing halo appearing above an area, icons shaking, and text labels slowly appearing;

[0130] Sound effects prompts: Play soft prompts, play voice guidance such as "You can view the Five-Colored Pool scenic spot";

[0131] Path guidance type: Displays a dashed trajectory in the user's view to guide the user to rotate the view or click to enter;

[0132] Combined event: Visual icon + sound effect + path guidance are triggered together.

[0133] The event is presented in a non-mandatory manner, without interfering with the current operator's actions, serving only as a subtle prompt to attract attention.

[0134] Step S405: Record user behavior feedback for subsequent system learning and strategy optimization. The recorded behavioral data includes:

[0135] Has the user entered the focus area?

[0136] The time interval between the appearance of the prompt and the entry;

[0137] User dwell time within the area;

[0138] Whether to click on resources within the area, and whether to trigger subsequent interaction processes;

[0139] A list of focus areas that are not being responded to from the current perspective.

[0140] After being structured, the above data is stored in an interaction log database for use by the guidance strategy module for periodic popularity analysis, interaction behavior modeling, and recommendation logic iteration. For example, the focus threshold can be adjusted based on popularity distribution, the intensity of visual guidance can be adjusted based on click-through rate, and the voice content can be adjusted based on failure prompt rate.

[0141] In a specific example, in the Aba Prefecture ecotourism themed exhibition hall, when several visitors were discussing "The colors of the Five-Colored Pool are truly unique," "Have you seen the Five-Colored Pool in autumn?" and "Is it on the right?", the system, through voice aggregation analysis, identified "Five-Colored Pool" as the group's center of interest. Even though the guide was still introducing "Shuzheng Qunhai" and hadn't noticed this semantic focus, the system generated a transparent breathing halo above the corresponding 3D area of ​​the Five-Colored Pool, accompanied by a soft prompt, "New hotspot area ready." This icon slowly appeared at the edge of the guide's current field of vision and gradually moved towards the Five-Colored Pool, simulating a guiding clue. Once the guide noticed this prompt, they could independently control the interface to switch to that area, completing the guidance loop from group intention to active control response.

[0142] Step S50: Based on the interactive operations of the multiple users on the focus area, display the pre-loaded content on the large screen.

[0143] In a large-screen 3D digital display environment with multi-user collaboration, preloading the content of the focus area is only part of the system response optimization. The ultimate key is to quickly present the preloaded content to the large screen after the user makes a clear interactive operation. Since the content displayed on the large screen system needs to reflect the user's behavioral intent in real time, it is necessary to quickly call up the cached resources of the corresponding focus area and visualize it from the target perspective after the user issues a clear signal (such as a voice command, click, gesture, or view lingering). This process should ensure the rapid responsiveness of the presentation behavior, the natural and smooth transition of the display, and the compatibility with various input forms, thereby constructing a complete closed-loop interactive process of "semantic perception - content preparation - operation triggering - visual response".

[0144] In one optional implementation, step S50 specifically includes the following sub-steps:

[0145] Step S501: Monitor interactive events. The system listens to user input devices in real time, including microphones, touchscreens, mice, motion-sensing cameras, and laser pointers, identifies event types and trigger locations, and records the time of occurrence, trigger identifier, and input category. For example, voice input "view the Five-Colored Pool," or the user continuously focusing their gaze on a focal area for more than three seconds.

[0146] Step S502: Map the interactive event to the focus area set R_focus, and determine whether the event hits a pre-loaded area. Event mapping methods include: spatial coordinate mapping (e.g., whether the 3D coordinates corresponding to the click point fall within the target area grid range), semantic parsing mapping (e.g., matching keywords in speech with area labels), and spatial intersection determination of device pointing trajectory and grid number.

[0147] Step S503: Invoke the preloaded resources for the corresponding focus area. If the event hit area is a mesh area marked as "ready" within R_focus, the resource module is loaded directly from the local cache index. The resource module includes a 3D structural model, texture maps, a multimedia playback controller, interactive hotspot binding information, etc. The loading method is based on WebGL or UnrealPixelStreaming engine, and the rendering preparation is completed in an asynchronous non-blocking manner.

[0148] Step S504: Perform main view switching and area focus display. Based on the center coordinates of the hit area, the system adjusts the position and orientation parameters of the 3D camera, allowing the user's viewpoint to naturally transition to the target area. The system then uses animation to advance the viewpoint, zoom in, or rotate the camera to center the target area on the screen. Subsequently, the area content is presented, including a floating text annotation window, an embedded video frame, an image carousel component, and background sound effects.

[0149] Step S505: Record and evaluate the display effect. The system records data such as the start time of the area display, duration, frequency of user interaction (e.g., clicking on interaction points, pausing the video, sliding the carousel), and subsequent feedback behavior (e.g., whether to further request the relevant area), and updates the content popularity model in real time for subsequent guidance strategy adjustments.

[0150] The advantages of this method lie in its ability to improve response speed not only through preloading the focus area but also through compatible listening and precise mapping of various user actions, achieving a low-latency, highly immersive display process for content triggered by multiple inputs. Adaptive view transitions and seamless resource loading prevent user experience interruptions due to waiting or accidental touches, giving large-screen interactive displays stronger responsiveness, adaptability, and context awareness, making them particularly suitable for group-oriented digital exhibition scenarios.

[0151] For example, in the Aba Prefecture Ecotourism Theme Exhibition Hall, a guide is introducing "Shuzheng Lakes" to a large screen, while several visitors are discussing the "Five-Colored Pool" attraction via voice chat. The system identifies "Five-Colored Pool" as the focus area and pre-loads resources accordingly. When the guide points towards "Five-Colored Pool" and says, "Look at the colors here," the system immediately recognizes this as a hit event, calls the cached resources, smoothly zooms in on the 3D model of Five-Colored Pool on the large screen, and simultaneously plays an explanatory video on "The Principle of Five-Colored Pool's Autumn Color Change," displaying accompanying text and images. This achieves a natural unity of content focus, interactive response, and contextual interpretation.

[0152] For any module structures not specifically defined in this invention, the existing technical descriptions shall prevail. The prior art mentioned in the foregoing background and specific embodiments sections can be considered part of this invention and used to understand the meaning of certain technical features or parameters.

Claims

1. A method for optimizing web interaction latency, characterized in that, The web is used to display 3D digital scenes on large-screen devices, where the large screen displays content simultaneously to multiple users. The method includes the following steps: The target area in the three-dimensional digital space of the Web terminal is divided into spatial grids, and a spatial index structure is constructed. The spatial index structure is used to describe the spatial position relationship and visual attributes of each grid area in the three-dimensional scene. Collect voice communication data between the multiple users, perform semantic parsing on the voice data to extract keywords, intent expressions and contextual semantic information, and construct a semantic vector model of the current focus of the multiple users; The semantic vector model is semantically matched with the label information of each grid region in the three-dimensional digital space to determine the target regions that the multiple users may be interested in later, as a set of potential focus regions. Content resources are preloaded onto the set of potential focus areas, and guided micro-interaction events are generated on the potential focus areas to induce the perspectives or interaction behaviors of the multiple users to converge on the preloaded areas. Based on the interactive operations of the multiple users on the focus area, the preloaded content is displayed on the large screen.

2. The Web interaction latency optimization method according to claim 1, characterized in that, The spatial grid is divided into regular cubes or deformed tetrahedral structures, and a spatial index structure is constructed that includes node positions, adjacent topological relationships, visible range, and occlusion attributes.

3. The Web interaction latency optimization method according to claim 1, characterized in that, The voice communication data is collected by an omnidirectional microphone array installed at the top of the display area or on the desktop, and background noise and non-voice signals are filtered out by a noise suppression algorithm.

4. The Web interaction latency optimization method according to claim 1, characterized in that, The process of constructing the semantic vector model includes inputting the text transcribed from speech recognition into a natural language processing model based on a deep neural network, and outputting a semantic embedding vector of fixed dimensions. The natural language processing model is at least one of Word2Vec, GloVe, or BERT.

5. The Web interaction latency optimization method according to claim 1, characterized in that, The tag information is a set of semantic tags preset for each grid area when constructing a 3D scene. The tag set includes scenic spot name, historical background, cultural attributes, geographical features or ecological theme text information.

6. The Web interaction latency optimization method according to claim 1, characterized in that, The semantic matching uses a cosine similarity algorithm to calculate the similarity between the semantic vector and the label vector, and sets a matching threshold or selects the top N regions with the highest similarity to form a set of potential focus regions.

7. The Web interaction latency optimization method according to claim 1, characterized in that, The preloading of the content resources includes 3D model files, texture layers, multimedia materials and interactive scripts, and adopts an asynchronous loading mechanism and a layered loading strategy, prioritizing the loading of low-resolution resources.

8. The Web interaction latency optimization method according to claim 1, characterized in that, The guided micro-interaction events include generating dynamic icons, glowing boundaries, indicating paths, or voice prompts in the focus area to guide the user's viewpoint toward the target area or trigger further interaction.

9. The Web interaction latency optimization method according to claim 1, characterized in that, The user's interactive operations include voice commands, mouse clicks, gestures, prolonged viewing time, or laser pointer trajectory input. These operations are identified in real time by the event listening module and mapped to the spatial index structure.

10. The Web interaction latency optimization method according to claim 1, characterized in that, After displaying the pre-loaded content on the large screen, the system also records and analyzes the user's actual response behavior to update the regional heat map, optimize resource scheduling order, and adjust the guidance event generation strategy.