Blogger live video AR glasses scenery labeling method and labeling system
By collecting and analyzing interactive data in real time, combining 3D spatial topology maps and NeRF algorithms, dynamically adjusting the content of AR glasses labeling, solving the problem of mismatching the content of the labeling in the existing technology and the needs of the audience, achieving more efficient audience interest and emotion capture, and improving the real-time and interactiveness of live broadcasts.
Patent Information
- Application Number
- CN202510769103.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing blogger live video AR glasses scene text labeling system cannot fully explore the value of interactive data, resulting in the labeling content that does not match the audience's needs, cannot respond to changes in the live broadcast scene in a timely manner, ignore the audience's group behavior and emotional tendencies, and affect the viewing experience.
By collecting video streams and interactive data streams in real time, using the Apache Flink stream processing engine to analyze barrage word frequency and emotional classification, combining IMU units and geolocation to construct a 3D spatial topology map, using NeRF algorithm to generate a scene enhancement labeling layer, and adjust the labeling visibility through an adaptive transparency algorithm and dynamically adjust the labeling content.
Real-time dynamic annotation of AR glasses landscape content, accurately capture the changes in audience interests and emotions, improve the real-time and accuracy of live broadcasts, meet the personalized needs of the audience, and enhance the viewing experience and interactivity.
Smart Images

Figure CN120281967A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AR glasses, and in particular to a method and system for annotating scenes and texts of AR glasses used in blogger live broadcast videos. Background Art
[0002] In today's digital age, the live broadcast industry is booming. Travel live broadcast, as a new form of live broadcast, has attracted a large number of viewers with its unique immersive experience. Bloggers can provide viewers with a more intuitive and vivid display of travel scenes by using AR glasses to broadcast travel live. However, although the existing blogger live broadcast video AR glasses landscape annotation system has achieved multimodal interaction and dynamic scene reconstruction to a certain extent, there are still some problems that need to be solved, which limit the quality of live broadcast and the audience's viewing experience.
[0003] First of all, at present, the interactive data such as bullet screens and likes in the live broadcast room contain rich real-time interest information of the audience. However, the existing technology has failed to fully tap the value of these interactive data, and only displays them as simple information without in-depth analysis of the changing trends of audience interests reflected behind them. For example, during a travel live broadcast, the audience may ask detailed questions about a specific attraction through bullet screens, or show strong interest in a certain attraction, but the system cannot capture this information in time and make corresponding adjustments, resulting in the mismatch between the scenery text annotation content of the AR glasses and the actual needs of the audience, and the inability to accurately push the information of interest to the audience.
[0004] Secondly, the existing technology has a significant lag in updating the annotated content. The system cannot dynamically adjust the annotated content according to the real-time interactive data in the live broadcast room, and often annotates according to preset rules and content. This results in the inability of the annotated content to keep up with the audience's interests or new hot topics during the live broadcast, making the information push out of touch with the audience's actual needs. For example, when a suddenly popular tourist attraction or a topic that sparks heated discussion among the audience appears in the live broadcast, the system cannot quickly integrate this information into the annotated content, affecting the audience's viewing experience.
[0005] Finally, existing technologies mainly focus on the analysis of preferences of AR glasses wearers, while ignoring the real-time behavior patterns of audience groups in live tourism. In live broadcast scenarios, the behavior of audience groups often has group characteristics and trends, such as the emergence of hot topics and the shift of collective interests. Existing technologies lack the ability to monitor and analyze the behavior of audience groups, and are unable to capture changes in these group behaviors in a timely manner, and are therefore unable to optimize and adjust the annotation content based on group behavior. For example, when most viewers are interested in a specific type of attraction, the system cannot automatically adjust the annotation focus to highlight this type of attraction, resulting in the live broadcast content being unable to meet the needs of the group.
[0006] In addition, the emotional tendencies of the audience during the live broadcast, such as excitement, confusion, and affection, are crucial for enhancing the interactivity and attractiveness of the live broadcast. However, the existing technologies lack the effective ability to identify the emotional tendencies of the audience and cannot provide corresponding emotional responses based on the emotional feedback of the audience. For example, when the audience shows excitement and affection for a certain scenic spot, the system cannot further enhance this emotional experience by adjusting the annotation content or display method; when the audience has doubts about a certain content, the system cannot give a clear explanation and annotation in a timely manner, affecting the viewing satisfaction of the audience. Summary of the Invention
[0007] The purpose of the present invention is to provide a method and a system for annotating the scenery text of the blogger's live video AR glasses. By collecting the video stream and interactive data stream in real time, the dynamic annotation of the scenery text content of the AR glasses is realized, which can timely respond to the changes in the live broadcast scene and the interactive needs of the audience, and improve the real-time performance and accuracy of the live broadcast annotation, so as to solve at least one of the above-mentioned problems of the existing technologies.
[0008] In the first aspect, the present invention provides a method for annotating the scenery text of the blogger's live video AR glasses. The method specifically includes: Collecting the video stream of the tourism live broadcast scene in real time through the built-in camera and built-in sensor of the AR glasses, and obtaining the interactive data stream in the live broadcast room in real time through the Apache Flink stream processing engine; Based on the interactive data stream in the live broadcast room, obtaining the set of hot topics and the emotional classification result by analyzing the bullet screen word frequency, live broadcast time period, and bullet screen text; Based on the video stream of the live broadcast scene, dynamically adjusting the knowledge graph retrieval strategy of the tourism live broadcast scene through the set of hot topics and the emotional classification result, and obtaining the dynamic retrieval result; Obtaining the blogger's head pose data based on the IMU unit of the AR glasses, and combining the geographic location information to construct a 3D spatial topology map of the tourism live broadcast scene through the SLAM algorithm; Performing spatial registration on the dynamic retrieval result and the 3D spatial topology map, and combining to generate a scene enhancement annotation layer through the NeRF algorithm; Overlaying and displaying the video stream and the scene enhancement annotation layer, and dynamically adjusting the annotation visibility by using the adaptive transparency algorithm.
[0009] In the second aspect, the present invention provides a system for annotating the scenery text of the blogger's live video AR glasses. The system specifically includes: The first annotation module is used to collect the video stream of the tourism live broadcast scene in real time through the built-in camera and built-in sensor of the AR glasses, and obtain the interactive data stream in the live broadcast room in real time through the Apache Flink stream processing engine; The second annotation module is used to obtain a set of hot topics and sentiment classification results based on the live stream interaction data stream by analyzing the bullet screen word frequency, live broadcast time period, and bullet screen text; The third annotation module is used to dynamically adjust the knowledge graph retrieval strategy of the tourism live broadcast scene based on the video stream of the live broadcast scene through the set of hot topics and sentiment classification results, and obtain dynamic retrieval results; The fourth annotation module is used to obtain the blogger's head pose data based on the IMU unit of the AR glasses, combine the geographic location information, and construct a 3D spatial topology map of the tourism live broadcast scene through the SLAM algorithm; The fifth annotation module is used to perform spatial registration on the dynamic retrieval results and the 3D spatial topology map, and combine to generate a scene enhancement annotation layer through the NeRF algorithm; The sixth annotation module is used to overlay and display the video stream and the scene enhancement annotation layer, and dynamically adjust the annotation visibility using an adaptive transparency algorithm.
[0010] In a third aspect, the present invention provides a computer device, including: a memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, it implements the blogger live video AR glasses scene text annotation method as described in any one of the above methods.
[0011] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the blogger live video AR glasses scene text annotation method as described in any one of the above methods.
[0012] Compared with the prior art, the present invention has at least one of the following technical effects: 1. By real-time collecting the video stream and interaction data stream, dynamic annotation of the scene text of the AR glasses is realized, which can timely respond to the changes in the live broadcast scene and the interaction needs of the audience, and improve the timeliness and accuracy of live broadcast annotation.
[0013] 2. Deeply analyze the interaction data such as bullet screens and likes in the live broadcast room, accurately capture the real-time interest changes of the audience, provide a strong basis for optimizing the annotation content, and make the annotation content more in line with the needs of the audience.
[0014] 3. Dynamically adjust the annotation content according to the interaction data, ensuring that the interest points and hot topics of the audience can be timely reflected during the live broadcast, avoiding the disconnection between information push and the needs of the audience, and improving the timeliness and pertinence of the live broadcast.
[0015] 4. Pay attention to the real-time behavior patterns of the audience group, timely capture group behaviors such as the emergence of hot topics, optimize the annotation content according to the group behavior, meet the common needs of the group, and improve the attractiveness and influence of the live broadcast.
[0016] 5. Identifying the emotional tendency of the audience and providing emotional responses according to the emotional feedback enhances the emotional interaction between the audience and the travel live broadcast, and improves the viewing satisfaction and loyalty of the audience.
[0017] 6. Splitting the bullet screen data by time window and analyzing the word frequency and text emotion can accurately capture the hot topics and the emotional tendency of the audience during different live broadcast periods, providing a reliable basis for subsequent annotation optimization.
[0018] 7. Dynamically adjusting the knowledge graph retrieval strategy by combining the hot topics and the results of emotion classification can screen out scene entities that better meet the interests and emotional needs of the audience, making the annotation content more targeted and attractive.
[0019] 8. Constructing a 3D spatial topology map based on the IMU unit data and geographical location information can accurately restore the spatial structure of the travel live broadcast scene, providing a more intuitive spatial positioning basis for subsequent annotation.
[0020] 9. Generating a scene enhancement annotation layer through spatial registration and the NeRF algorithm can fuse the dynamic retrieval results with the 3D spatial information, realizing the deep combination of the annotation and the scene, and enhancing the visual effect of the annotation.
[0021] 10. Adopting an adaptive transparency algorithm to dynamically adjust the annotation visibility, and comprehensively determining the transparency according to the blogger's fixation point, head movement and the importance of the annotation content, making the annotation display more in line with the blogger's viewing habits and improving the viewing experience.
[0022] 11. Combining the audience portrait data and the real-time feedback data, calculating the display weight of the annotation elements through the semantic similarity function and generating a rendering strategy matrix can realize personalized annotation display, meet the needs of different audiences, and improve the interactivity and audience satisfaction of the live broadcast. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1 is a schematic flowchart of a method for annotating the scene text of a blogger's live video with an AR glasses; Figure 2 is a schematic structural diagram of a system for annotating the scene text of a blogger's live video with an AR glasses; Figure 3 is a schematic structural diagram of a computer device. Detailed implementation manners
[0025] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0026] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0027] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.
[0029] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0030] The reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0031] In the embodiments of the present application, the execution subject of the process includes a terminal device. The terminal device includes, but is not limited to, devices such as servers, computers, smartphones, and tablets that can execute the methods disclosed in the present application. Figure 1 The following is a schematic flowchart showing the method for Jingwen annotation of a blogger's live video AR glasses disclosed in an embodiment of the present invention: S101, Real-time collect the video stream of the travel live broadcast scene through the built-in camera and built-in sensors of the AR glasses, and obtain the interactive data stream in the live broadcast room in real time through the Apache Flink stream processing engine.
[0032] In this embodiment, through standardization processing and timestamp synchronization, the consistency and accuracy of the video stream and the interactive data stream are ensured, and the reliability of the analysis results is improved. The Apache Flink stream processing engine has good scalability and can easily handle the processing requirements of high-concurrency data streams, providing strong support for the dynamic annotation and personalized recommendation of travel live broadcasts.
[0033] S102, Based on the interactive data stream in the live broadcast room, obtain the set of hot topics and the sentiment classification results by analyzing the bullet screen word frequency, live broadcast period, and bullet screen text.
[0034] In this embodiment, the interactive data is sliced according to the live broadcast period for subsequent analysis of the audience interaction in different time periods. The bullet screen text in each data slice is segmented, and the continuous text is cut into meaningful lexical units. A dictionary-based segmentation method or a statistics-based segmentation method can be used. Count the frequency of each lexical unit in the bullet screen text to form a word frequency statistical table. Focus on the words with higher frequencies, which may represent the hot topics currently concerned by the audience.
[0035] According to the word frequency statistical results, identify the keywords that may represent hot topics. The keywords can be screened out by setting a word frequency threshold or using a topic extraction algorithm (such as the TF-IDF algorithm). Aggregate the relevant keywords together to form a set of hot topics. Analyze the heat change of hot topics in different live broadcast periods, and associate and store the hot topics with the corresponding live broadcast periods for subsequent dynamic adjustment of the annotation content according to the live broadcast period.
[0036] Collect existing basic sentiment lexicons, including positive sentiment words, negative sentiment words, and neutral sentiment words. These lexicons can serve as the basis for sentiment classification. For the field of travel live streaming, expand the sentiment lexicon. Match sentiment words in the barrage text of each data slice, and count the occurrences of positive sentiment words, negative sentiment words, and neutral sentiment words. Determine the overall sentiment tendency of the barrage text based on the occurrence times and sentiment tendencies of the sentiment words. Count the number of barrage texts with different sentiment tendencies in each data slice to form the sentiment classification statistical results. Analyze the changing trend of the sentiment classification results over time to understand the emotional fluctuations of the audience during the live stream.
[0037] In this embodiment, it is possible to analyze the interactive data in the live streaming room in real time, timely capture hot topics and changes in the audience's emotions, and provide support for dynamically adjusting the annotation content.
[0038] S103, based on the video stream of the live streaming scenario, dynamically adjust the knowledge graph retrieval strategy for the travel live streaming scenario through the hot topic set and the sentiment classification results to obtain dynamic retrieval results.
[0039] In this embodiment, image recognition technology (such as an object detection algorithm based on deep learning) is used to identify the scene elements in the video frame, including scenic spot buildings, natural landscapes, people, signboards, etc. Combining natural language processing technology, semantic understanding is carried out on the identified scene elements, and corresponding semantic labels are assigned. Associate the hot topic set obtained by analyzing the interactive data stream in the live streaming room with the scene element recognition results. Analyze the occurrence frequency and heat changes of hot topics at different time periods of the live stream, and combine the appearance time and position information of the scene elements in the video stream to further refine the association degree between the hot topics and the scene elements.
[0040] According to the sentiment classification results, assign corresponding sentiment weights to the scene elements. For example, if the audience shows positive emotions (such as excitement, love) towards a certain scene element (such as a beautiful seaside scenery), a higher positive sentiment weight is assigned to this scene element; if the audience shows negative emotions (such as dissatisfaction, boredom) towards a certain scene element (such as the crowded people in the scenic area), a lower sentiment weight or a negative sentiment weight is assigned to this scene element.
[0041] Consider the change of sentiment tendency over time and dynamically adjust the sentiment weights of the scene elements. For example, as the live stream progresses, the audience's interest in a certain scene element that was originally positive gradually decreases, and the sentiment weight also decreases accordingly; conversely, if a certain scene element suddenly triggers enthusiastic discussions and positive sentiment feedback from the audience, its sentiment weight is quickly increased.
[0042] Before the live broadcast starts, set the initial knowledge graph retrieval strategy according to the theme and common scenarios of the travel live broadcast. The knowledge graph contains rich travel-related information, such as scenic spot introductions, historical cultures, travel guides, surrounding delicacies, etc. The initial retrieval strategy can be based on common travel information needs, such as sorting and retrieving according to the popularity and heat of scenic spots.
[0043] According to the set of hot topics, dynamically adjust the retrieval direction of the knowledge graph. For example, when the hot topic focuses on the historical culture of a specific scenic spot, place the retrieval focus on the knowledge graph nodes related to the historical culture of that scenic spot, such as the construction history of the scenic spot, important historical events, and relevant historical figures.
[0044] Combine the emotional weights of scene elements to optimize the sorting of the knowledge graph retrieval results. For scene elements with higher emotional weights, preferentially retrieve detailed information related to them and increase their display priority in the retrieval results.
[0045] As the live broadcast progresses, continuously adjust the knowledge graph retrieval strategy in real time according to new hot topics and emotion classification results. For example, when a new hot topic appears in the live broadcast or the emotional tendency of the audience changes significantly, immediately re-evaluate the association relationship and emotional weight of the scene elements, and update the retrieval strategy accordingly to ensure that the retrieval results always match the real-time interests and needs of the audience.
[0046] According to the dynamically adjusted knowledge graph retrieval strategy, retrieve information related to the current live broadcast scene, hot topics, and audience emotional tendency from the knowledge graph. This information includes but is not limited to detailed introductions of scenic spots, historical and cultural backgrounds, travel guide suggestions, surrounding food recommendations, relevant stories and legends, etc. Integrate and screen the retrieved information, remove duplicate, irrelevant, or low-quality information, and retain the most valuable content that best meets the needs of the audience. Organize the integrated and screened information according to a certain logic and structure to generate dynamic retrieval results. The dynamic retrieval results can be designed according to different display forms, such as text descriptions, picture displays, video clips, audio explanations, etc., to meet the diverse viewing needs of the audience.
[0047] In this embodiment, obtaining dynamic retrieval results by dynamically adjusting the knowledge graph retrieval strategy provides strong support for the scene text annotation and content display of travel live broadcasts, which helps to improve the quality of travel live broadcasts and the viewing satisfaction of the audience. The dynamic retrieval results can be combined with the live broadcast scene and the display function of AR glasses to provide a more immersive travel live broadcast experience for the audience.
[0048] S104, obtain the blogger's head pose data based on the IMU unit of the AR glasses, and combine the geographic location information to construct a 3D spatial topology map of the travel live broadcast scene through the SLAM algorithm.
[0049] In this embodiment, the IMU (Inertial Measurement Unit) built into the AR glasses includes sensors such as an accelerometer, a gyroscope, and a magnetometer. The accelerometer is used to measure the acceleration of an object in three axial directions, the gyroscope is used to measure the angular velocity of an object around three axes, and the magnetometer is used to measure the Earth's magnetic field strength to assist in determining the direction. The IMU unit is set to collect data at a relatively high frequency (e.g., 100 Hz) to ensure that the minute movements of the blogger's head can be captured in real time. The collected data is recorded in a specific format, such as fields including a timestamp, acceleration values (in the X, Y, and Z directions), angular velocity values (in the X, Y, and Z directions), and magnetic field strength values (in the X, Y, and Z directions). Preprocessing is performed on the collected head pose data, including operations such as denoising and filtering.
[0050] High-precision geolocation technologies are adopted, such as the Global Positioning System (GPS) combined with the Assisted Global Satellite Positioning System (A - GPS), or indoor positioning technologies such as Wi - Fi positioning and Bluetooth beacon positioning (selecting the appropriate positioning method according to different tourism live broadcast scenarios). Through the positioning module integrated on the AR glasses, the geographical location information of the blogger, including longitude, latitude, altitude, etc., is obtained in real time. The positioning data also records a timestamp for time synchronization with the head pose data. To improve the positioning accuracy, a fusion scheme of multiple positioning technologies can be adopted.
[0051] Since there may be a deviation in the acquisition clocks of the head pose data and the geolocation data, time synchronization is required. Methods such as hardware clock synchronization or software timestamp alignment are used to ensure that the head pose data and the geolocation data accurately correspond on the time axis. The time synchronization error is controlled within a small range (e.g., millisecond level) to ensure the accuracy and consistency of the data during subsequent SLAM algorithm processing.
[0052] The preprocessed head pose data and geolocation data are fused. Loose coupling or tight coupling fusion strategies are adopted. The loose coupling method is to use the head pose data and the geolocation data respectively for independent pose estimation and then fuse the results; the tight coupling method is to directly input the head pose data and the geolocation data into the same state estimator for joint optimization. Through data fusion, the high-frequency dynamic characteristics of the head pose data and the absolute position information of the geolocation data can be fully utilized to improve the accuracy and robustness of pose estimation.
[0053] According to the characteristics of the tourism live broadcast scenario, select an appropriate SLAM algorithm, such as a vision-based SLAM algorithm (VSLAM) combined with head pose data, or use a laser SLAM algorithm (if the AR glasses are equipped with a lidar sensor). At the beginning of the live broadcast, initialize the SLAM system. The initialization process includes selecting an appropriate reference frame (usually the first frame image at the start of the live broadcast), estimating the initial camera pose (combining geolocation information as the initial reference), and initializing the coordinate system and key parameters of the map.
[0054] Extract feature points from the real-time video frames captured by the built-in camera of the AR glasses. Commonly used feature extraction algorithms include SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), and ORB (Oriented FAST and Rotated BRIEF), etc. The ORB algorithm has the advantages of fast calculation speed and good real-time performance, and is suitable for use in scenarios with high real-time requirements such as tourism live broadcasts. Match the feature points extracted from the current frame with those of the previous frame to determine the relative motion relationship of the camera at different times. Use brute force matching or fast approximate nearest neighbor (FLANN) matching algorithms for feature matching, and remove incorrect matching point pairs through the Random Sample Consensus (RANSAC) algorithm to improve the accuracy of the matching.
[0055] Combine the head pose data and the results of feature matching to estimate the current pose of the camera. The head pose data provides the motion increment information of the camera, and the feature matching provides the relative pose constraints between the cameras. Optimally estimate the pose of the camera through an Extended Kalman Filter (EKF) or a non-linear optimization algorithm (such as the graph optimization algorithm g2o) to obtain more accurate pose information. According to the estimated camera pose, add the extracted feature points to the map to construct a 3D spatial point cloud map. At the same time, filter and optimize the point cloud in the map to remove redundant point cloud data and improve the quality and storage efficiency of the map. For example, use a voxel grid filtering algorithm to downsample the point cloud to reduce the number of point clouds.
[0056] During the live broadcast, continuously detect whether the camera has returned to a previously passed position, that is, perform loop detection. Loop detection can correct the pose errors accumulated during the SLAM process and improve the global consistency of the map. Use the Bag of Words model or deep learning methods for loop detection, and judge whether a loop has occurred by comparing the feature descriptors of the current frame and historical frames. When a loop is detected, use the loop information to optimize the global map and camera pose. Jointly optimize the pose constraints and loop constraints during the entire SLAM process through a graph optimization algorithm to further reduce the pose error and obtain a more accurate 3D spatial topological map.
[0057] In this embodiment, by combining the IMU unit data and the geolocation information and adopting an advanced SLAM algorithm, a high-precision 3D spatial topological map can be constructed, accurately reflecting the spatial structure and layout of the tourism live broadcast scene. The constructed 3D spatial topological map provides a basis for scenic text annotation and virtual content fusion, and can provide a more intuitive, vivid and immersive tourism live broadcast experience for the audience, meeting the audience's needs for tourism information acquisition and visual enjoyment.
[0058] S105, spatially register the dynamic retrieval results and the 3D spatial topological map, and combine to generate a scene enhancement annotation layer through the NeRF algorithm.
[0059] In this embodiment, a detailed analysis is performed on the dynamic retrieval results obtained after dynamically adjusting the tourism live broadcast scene knowledge graph retrieval strategy based on the live broadcast scene video stream, the set of hot topics, and the sentiment classification results. The dynamic retrieval results usually contain various information related to the current live broadcast scene, such as detailed introductions of scenic spots, historical and cultural backgrounds, content related to real-time hot topics, etc., and this information exists in a structured or semi-structured form. The parsed information is classified and sorted according to certain rules, for example, divided according to dimensions such as scenic spots and topics, so as to be spatially registered with the 3D spatial topological map later.
[0060] Feature extraction is performed on the 3D spatial topological map of the tourism live broadcast scene constructed by the SLAM algorithm. The 3D spatial topological map contains the spatial structure information of the live broadcast scene, such as the positions and shapes of buildings, the distribution of scenic spots, etc. Key feature points, feature surfaces or feature regions can be extracted from the map, such as the vertices of buildings, the center points of scenic spots, etc. as reference features for registration. At the same time, the map is gridded, dividing the map into multiple small grid units to facilitate subsequent matching and registration operations.
[0061] Adopt a feature-based registration strategy to match the information in the dynamic retrieval results with the features in the 3D spatial topological map. For example, for a specific scenic spot mentioned in the dynamic retrieval results, search for the corresponding feature region in the 3D spatial topological map. Matching can be performed by calculating the similarity between features, such as the distance between feature points, the shape similarity of feature surfaces, etc. When a matching feature is found, determine the spatial position of the scenic spot information in the dynamic retrieval results in the 3D spatial topological map.
[0062] In addition to feature - based registration, semantic associations are also considered for registration. Analyze the semantic information in the dynamic retrieval results, such as the type of scenic spots, historical and cultural backgrounds, etc., and associate it with the semantic information of different regions in the 3D spatial topological map. For example, if an ancient building scenic spot with a profound historical and cultural heritage is mentioned in the dynamic retrieval results, then search for the region related to history and culture in the 3D spatial topological map and register the information of this scenic spot to the corresponding location. By combining feature matching and semantic associations, the accuracy and reliability of registration are improved.
[0063] Then, perform initial registration by quickly matching some obvious feature information in the dynamic retrieval results with the corresponding features in the 3D spatial topological map to determine a rough registration position. For example, match the location of the landmark scenic spot mentioned in the dynamic retrieval results with the center point of this scenic spot in the 3D spatial topological map as the benchmark for initial registration. Based on the initial registration, perform fine registration. Adopt the idea of the Iterative Closest Point (ICP) algorithm to continuously adjust the relative position and attitude between the dynamic retrieval results and the 3D spatial topological map to minimize the feature matching error between the two. Through multiple iterations, gradually optimize the registration result until a satisfactory registration accuracy is achieved.
[0064] Visualize the registered dynamic retrieval results and the 3D spatial topological map, and verify the accuracy of registration by means of manual observation. Check whether the information in the dynamic retrieval results is accurately located at the corresponding position in the 3D spatial topological map and whether there are obvious deviations or errors. If an error is found in the registration result, analyze the error. It may be caused by inaccurate feature extraction, unreasonable parameter settings of the registration algorithm, etc. According to the results of error analysis, optimize and adjust the feature extraction method, registration algorithm parameters, etc., and perform the registration operation again until the registration result meets the requirements.
[0065] The NeRF (Neural Radiance Fields) algorithm is a neural network-based method for scene representation and rendering. It represents a scene as a continuous 5D function (taking spatial position and viewing direction as inputs and color and density as outputs), and uses a neural network to learn the geometric and appearance information of the scene. During training, a large number of scene images at different viewpoints and their corresponding camera parameters are input. The neural network learns from these data to predict the color and density at any position and viewpoint in the scene, thereby achieving high-quality scene rendering. In the tourism live broadcast scenario, the NeRF algorithm can be used to generate a more realistic and delicate scene enhancement annotation layer. The tourism live broadcast scenario usually contains rich landscapes and details, and traditional annotation methods may not be able to present these details well. The NeRF algorithm can learn the three-dimensional structure and appearance features of the scene based on the video stream data of the live broadcast scenario, generate a scene rendering result with high realism, and provide more abundant visual information for scene-text annotation.
[0066] Collect the video stream of the tourism live broadcast scenario from multiple different viewpoints through the built-in camera of the AR glasses. These viewpoints should cover as many parts of the live broadcast scenario as possible to provide rich scene information. During the collection process, record the camera parameters of each viewpoint, such as the position, orientation, and focal length of the camera. Perform operations such as denoising and color correction on the video frames to improve the image quality. Then, classify and organize the video frames according to the viewpoints to construct a training dataset. At the same time, normalize the camera parameters to meet the input requirements of the NeRF algorithm. Design a NeRF model architecture suitable for the tourism live broadcast scenario. Generally, the NeRF model consists of multi-layer perceptrons (MLPs) for learning the 5D function of the scene. The parameters such as the number of model layers and the number of neurons can be adjusted according to the characteristics of the tourism live broadcast scenario to improve the performance and training efficiency of the model. Input the preprocessed training dataset into the NeRF model for training. During training, use optimization algorithms such as stochastic gradient descent to continuously adjust the model parameters to minimize the error between the color and density predicted by the model and the real scene images. Through multiple iterative trainings until the model converges, a NeRF model that can accurately represent the tourism live broadcast scenario is obtained.
[0067] Fuse the annotation information in the spatially registered dynamic retrieval results with the scene rendering results generated by the NeRF model. The annotation information can include text information such as the name, introduction, historical and cultural background of scenic spots, as well as some simple graphical annotations, such as arrows, icons, etc. Display the text information at the corresponding positions in the scene rendering results with appropriate fonts and colors, and associate the graphical annotations with the objects in the scene to make them blend naturally with the scene. Optimize the visual effect of the enhanced annotation layer in the fused scene. Adjust parameters such as the transparency and color contrast of the annotation information so that it neither affects the realism of the scene nor fails to be clearly presented to the audience. At the same time, consider the layout and hierarchy of the annotation information to avoid overcrowding or overlapping of the annotation information and improve the readability and aesthetics of the annotation layer.
[0068] In this embodiment, by spatially registering the dynamic retrieval results with the 3D spatial topological map, the scenic text annotation information can be accurately located at specific positions in the live broadcast scene, avoiding the problem of mismatch between the annotation information and the actual scene and improving the accuracy of the annotation. Combining with the NeRF algorithm to generate the enhanced annotation layer for the scene, taking advantage of the high-fidelity rendering ability of the NeRF algorithm, provides a richer and more delicate visual background for the scenic text annotation, enabling the audience to obtain a more immersive experience when watching the live broadcast. It can process the data during the live broadcast in real time, and dynamically adjust the dynamic retrieval results and the enhanced annotation layer for the scene according to the interaction data in the live broadcast room and the changes in the live broadcast scene, meeting the real-time and dynamic requirements of tourism live broadcasts. The combination of accurate annotation and realistic scene rendering can better meet the audience's needs for tourism information, improve the audience's viewing satisfaction and participation, and bring better interaction effects to tourism live broadcasts.
[0069] S106, superimpose and display the video stream and the enhanced annotation layer for the scene, and dynamically adjust the annotation visibility using an adaptive transparency algorithm.
[0070] In this embodiment, the enhanced annotation layer for the scene contains annotation information generated based on the dynamic retrieval results and the 3D spatial topological map, such as text information like the name, introduction, historical and cultural background of scenic spots, as well as some graphical annotations, such as arrows, icons, etc. Perform format conversion and optimization processing on the enhanced annotation layer for the scene to make it match the format and size of the video frames. Display the video stream and the enhanced annotation layer for the scene in a layered superimposed manner. Use the video frame as the bottom layer image and the enhanced annotation layer for the scene as the upper layer. During the display process, ensure that the upper annotation layer can accurately cover the corresponding positions on the bottom video frame to achieve the precise correspondence between the annotation information and the actual scene.
[0071] Reasonably plan the display area according to the importance and type of annotation information. Place important annotation information, such as the core introduction of scenic spots, in prominent positions on the video frame, such as the central area of the picture or positions closely related to the scenic spots; for some auxiliary annotation information, such as arrow direction indicators, they can be placed at the edge of the picture or in positions that do not affect the display of the main scene. At the same time, consider the layout and spacing between annotation information to avoid overcrowding or overlapping of annotation information, which may affect the viewing experience of the audience.
[0072] During the overlay display process, ensure frame synchronization between the video frame and the scene enhancement annotation layer. That is, each video frame should be accurately overlaid with the corresponding annotation layer. Frame synchronization can be achieved through timestamps or other synchronization mechanisms to ensure that during playback, the annotation information can be accurately updated and displayed as the video frame switches. Use graphics rendering technology to render the overlaid image. During the rendering process, consider factors such as color fusion and lighting effects of the image to make the annotation information blend naturally with the video frame and form a unified visual picture. Finally, output the rendered image to the display device for the audience to watch.
[0073] Evaluate the audience's interest in the current annotation information by analyzing the interaction data in the live broadcast room, such as the number of bullet comments, likes, and comment content. If the audience shows a high level of interest in a certain annotation information, for example, a large number of audiences discuss the scenic spot corresponding to the annotation in the bullet comments, then appropriately reduce the transparency of the annotation information to make it more prominent and attract the audience's attention; on the contrary, if the audience's interest in a certain annotation information is low, then appropriately increase its transparency to reduce the interference to the audience's viewing of the main scene.
[0074] Consider the impact of the complexity of the tourism live broadcast scene on the visibility of annotations. In the case of a relatively simple scene with less information, the transparency of the annotation information can be appropriately reduced to make it more clearly visible; while in a complex scene with a large amount of information, such as a scene with dense scenic spots and a large number of people, the transparency of the annotation information can be appropriately increased to avoid the annotation information being too prominent and covering up the details of the main scene.
[0075] Record the viewing duration of the audience for each annotation information. If the audience has a relatively long viewing duration for a certain annotation information, it indicates that the audience is more concerned about this information. At this time, the transparency can be reduced to further enhance the display effect of the information; if the viewing duration is short, then increase the transparency to reduce the interference to the audience.
[0076] Assign corresponding weights to the above different adjustment factors. For example, the audience interest may be given a higher weight because it is a key factor directly affecting the audience's viewing experience; the scene complexity and viewing duration are assigned relatively lower weights according to the actual situation. Through reasonable weight assignment, comprehensively consider the impact of various factors on the transparency of annotation information.
[0077] Set a reasonable range for the transparency of the annotation information, for example, from 0% (completely opaque) to 100% (completely transparent). According to different adjustment factors and weights, calculate the target transparency value of each annotation information in the current situation, and ensure that this value is within the set transparency range.
[0078] During the live broadcast, real-time monitor information such as the interaction data in the live broadcast room, the scene complexity, and the viewing duration of the audience. According to the set adjustment strategy and weights, calculate the target transparency value of each annotation information in real time. For example, analyze the interaction data every certain period (such as 1 second), update the evaluation result of the audience's interest, and recalculate the transparency of the annotation information.
[0079] To avoid discomfort to the audience caused by sudden changes in the transparency of the annotation information, adjust the transparency in a smooth transition manner. For example, when the target transparency value changes, instead of immediately adjusting the transparency to the target value, gradually transition to the target value within a certain time interval (such as 0.5 seconds) to make the change of transparency more natural and smooth.
[0080] Through the interactive functions in the live broadcast room, such as bullet screens and comments, collect the feedback of the audience on the adjustment effect of the transparency of the annotation information. Understand whether the audience thinks that the display of the annotation information is clearer and more reasonable, and whether it has a positive impact on the viewing experience. Optimize and adjust the adaptive transparency algorithm according to the audience feedback and the actual adjustment effect.
[0081] In this embodiment, by accurately superimposing and displaying the video stream and the scene enhancement annotation layer, and using the adaptive transparency algorithm to dynamically adjust the annotation visibility, the annotation information can be more clearly and accurately displayed in front of the audience, improving the efficiency and accuracy of information transmission. Accurate and clear annotation information and reasonable transparency adjustment can attract the audience's attention, stimulate the audience's participation enthusiasm, promote the interaction between the audience and the anchor, and improve the interactivity and attractiveness of the live broadcast. It can automatically adjust the display method of the annotation information according to different tourism live broadcast scenarios and the characteristics of the audience group, with strong adaptability and flexibility, and can meet the needs of various tourism live broadcast scenarios.
[0082] In some embodiments, in the above step S102, based on the interactive data stream in the live broadcast room, by analyzing the bullet screen word frequency, the live broadcast period, and the bullet screen text, obtaining the set of hot topics and the sentiment classification result specifically includes: Extract the bullet screen data set from the interactive data stream in the live broadcast room, perform time window segmentation on the bullet screen data set, and divide it into several bullet screen data subsets according to the live broadcast period; Perform word frequency statistics on the subset of bullet screen data within each time window, calculate the heat weight value of each term, and obtain the word frequency vector statistical result; Based on the word frequency vector statistical result, generate a set of hot topics through an improved LDA topic model; Perform sentiment polarity classification on the bullet screen text of the subset of bullet screen data within each time window, use the RoBERTa-base model to calculate the sentiment score of each bullet screen, and obtain the sentiment classification result.
[0083] In this embodiment, parse the interactive data stream, extract the data fields related to the bullet screen, and construct a bullet screen data set. During the parsing process, it is necessary to clean the data to remove invalid data and noise data.
[0084] According to the characteristics of the live broadcast and business requirements, set an appropriate time window size. The size of the time window needs to comprehensively consider factors such as the rhythm of the live broadcast, the frequency of topic changes, and the limitations of computing resources. If the time window is set too small, it may lead to insufficient data volume within each window, making it difficult to accurately reflect hot topics and sentiment tendencies; if the time window is set too large, it may not be able to capture topic and sentiment changes in a timely manner.
[0085] According to the set time window size, split the bullet screen data set. Extract the bullet screen data within each time window to form an independent subset of bullet screen data. Select a word segmentation tool suitable for Chinese text, such as Jieba. Perform word segmentation on the subset of bullet screen data within each time window, splitting the continuous bullet screen text into individual terms. During the word segmentation process, stop word filtering is also performed. Stop words refer to words that appear frequently in the text but contribute little to the understanding of the text semantics, such as "de", "le", "shi", etc. By removing stop words, the data volume can be reduced, improving the accuracy and efficiency of subsequent word frequency statistics.
[0086] For the subset of bullet screen data within each time window, count the number of occurrences of each term. A data structure such as a hash table can be used to record the word frequency information, with the term as the key and the number of occurrences as the value. Store the word frequency statistical results within each time window for subsequent analysis and processing. The results can be stored in a database or saved in the form of a file to ensure data traceability and queryability. When calculating the heat weight value of a term, in addition to word frequency, other factors can also be considered, such as the position of the term in the bullet screen text and whether it is a newly emerging term.
[0087] Considering multiple factors, calculate the popularity weight value for each term. The weighted average method can be used. According to the importance of different factors, assign corresponding weights, and then sum the scores of each factor after weighting to obtain the final popularity weight value of the term. Combine all terms and their corresponding popularity weight values within each time window to form the term frequency vector statistical result. This vector can intuitively reflect the importance of each term within the time window, providing a data basis for subsequent hot topic generation.
[0088] The LDA (Latent Dirichlet Allocation) topic model is a commonly used text topic modeling method. It assumes that a document is generated by a mixture of multiple topics, and each topic is composed of the probability distributions of multiple terms. Through learning a large number of documents, the LDA model can discover the latent topic structure in the documents. Traditional LDA models may have some limitations when dealing with live barrage data, such as poor performance in processing short texts and inability to fully consider time factors. Therefore, it is necessary to improve the LDA model. The improvement directions can include introducing time dimension information to enable the model to capture the changes of topics over time, optimizing the method for determining the number of topics to improve the accuracy of topic mining, etc.
[0089] Convert the term frequency vector statistical result within each time window into a format suitable for LDA model training. Usually, it is necessary to convert the data into the form of a document-term matrix, where each row represents a time window (which can be regarded as a document), each column represents a term, and the elements in the matrix represent the popularity weight value of the term within the time window. Determine some key parameters of the LDA model, such as the number of topics K, hyperparameters α and β, etc. The number of topics K can be determined by some heuristic methods or methods based on cross-validation; the hyperparameters α and β can be set according to experience or experiments, and they affect the distributions of topics and terms.
[0090] Use the prepared training data to train the improved LDA topic model. During the training process, the model will continuously adjust the distribution parameters of topics and terms to enable the model to better fit the training data. After training, extract each topic from the model. Each topic consists of a set of terms with relatively high probabilities, and these terms can reflect the core content of the topic. According to the meaning of the terms and context information, interpret and name each topic to obtain the set of hot topics.
[0091] RoBERTa (Robustly Optimized BERT Pretraining Approach) is a pre-trained language model based on the Transformer architecture. It has made a series of improvements on the basis of BERT, such as using a larger batch size, longer training sequences, etc., thus improving the performance and generalization ability of the model. The RoBERTa-base model is a basic version in the RoBERTa series, with a smaller model size and faster inference speed, suitable for scenarios with high real-time requirements. Compared with traditional sentiment classification methods, the RoBERTa-base model can better understand the semantic information of the text and capture the sentiment tendency in the text. It learns rich language knowledge and semantic representations through pre-training on large-scale text data, and can accurately process various complex emotional expressions.
[0092] Collect a certain amount of live barrage text data and perform sentiment annotation. Sentiment annotation is usually divided into three categories: positive, negative, and neutral. The artificial annotation method can be used to annotate the barrage text, and then methods such as consistency checking are used to ensure the accuracy of the annotation results. Divide the annotated dataset into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used to adjust the hyperparameters of the model and prevent overfitting, and the test set is used to evaluate the final performance of the model.
[0093] Fine-tune the RoBERTa-base model using the prepared training set. During fine-tuning, replace the output layer of the model with a classifier suitable for the sentiment classification task (such as a softmax classifier), and adjust the parameters of the model so that the model can achieve better performance in the sentiment classification task. For each piece of barrage text in the subset of barrage data within each time window, input it into the fine-tuned RoBERTa-base model, and the model will output the probability values that the barrage text belongs to the three sentiment categories of positive, negative, and neutral. Take the probability value of positive sentiment as the sentiment score of this barrage (ranging from 0 to 1, the higher the score, the more positive the sentiment). Determine its sentiment polarity according to the sentiment score of each barrage. For example, a threshold can be set. When the positive sentiment score is greater than this threshold, determine that this barrage is of positive sentiment; when the negative sentiment score is greater than this threshold, determine it as negative sentiment; otherwise, determine it as neutral sentiment. Summarize the sentiment classification results of all barrages within each time window to obtain the sentiment classification result of this time window.
[0094] In this embodiment, an effective analysis of the hot topics and the audience's emotions during the live broadcast is realized, providing strong support for the optimization and personalized services of the travel live broadcast. According to the audience's sentiment tendency and the hot topics they are concerned about, provide personalized live content recommendations and annotation displays for the audience.
[0095] In some embodiments, in the above step S103, for the video stream based on the live broadcast scenario, the knowledge graph retrieval strategy of the travel live broadcast scenario is dynamically adjusted through the hot topic set and the sentiment classification result to obtain a dynamic retrieval result, which specifically includes: Extract the shooting direction angle and timestamp in the video stream of the live broadcast scenario, and generate a scene entity set through a target detection model. The scene entity set includes a number of scene entities and their bounding box coordinates and class labels; Based on the scene entity set, combine the hot topic set and the sentiment classification result to construct a dynamic retrieval weight vector; Based on the dynamic retrieval weight vector, screen out the first scene entity set from the pre-set knowledge graph retrieval strategies; According to the associated historical events and their association degrees of each first scene entity in the first scene entity set, and the overlap rate between the bounding box of each first scene entity and the visual attention area of the blogger, calculate the historical relevance score of each first scene entity in the first scene entity set; Based on the historical relevance score, screen out the second scene entity set from the first scene entity set to obtain a dynamic retrieval result.
[0096] In this embodiment, in the video stream data, each frame of image is accompanied by corresponding timestamp information. By parsing the encapsulation format of the video stream (such as MP4, FLV, etc.), the timestamp of each frame of image is extracted. The timestamp is used to identify the specific time position of the video frame during the live broadcast, providing a time reference for subsequent scene analysis and retrieval. The shooting direction angle information is obtained by using the sensors built in the AR glasses (such as gyroscopes, accelerometers, etc.). These sensors can monitor the orientation and angle changes of the camera in real time, and associate the obtained direction angle data with the video frames.
[0097] Select an object detection model suitable for live broadcasting scenarios, such as the YOLO (You Only Look Once) series of models or the Faster R-CNN model. These models have high detection speed and accuracy and can quickly identify various scene entities in real-time video streams. Input the extracted video frames into the object detection model for inference. The model analyzes the video frames, detects various objects in them, and generates bounding box coordinates and category labels for each object. For example, in a travel live broadcasting scenario, the model may detect scene entities such as buildings, people, and natural landscapes, and mark their locations (bounding box coordinates) and categories (such as "ancient buildings", "tourists", "mountains", etc.) in the video frame. Summarize all scene entities detected by the model and their corresponding bounding box coordinates and category labels to construct a scene entity set. Each scene entity exists in the set as a data entry, which contains information such as the entity's category, the coordinates of the upper left corner and the lower right corner of the bounding box, etc.
[0098] Conduct in-depth analysis on the generated hot topic set to understand the scene entity categories involved in each hot topic. For example, if the hot topic is "the characteristics of a certain ancient building", then the scene entity categories related to it may include "ancient building" and "architectural decoration". At the same time, count the frequency and popularity of each hot topic during the live broadcast to measure its importance to the current live broadcast content.
[0099] According to the sentiment classification results, analyze the audience's sentiment inclination towards different scene entities. For example, if the audience's barrage sentiment towards a certain scenic spot is mostly positive, then the weight of the scene entity related to the scenic spot in the retrieval can be appropriately increased. Count the number of barrages of different sentiment categories (positive, negative, neutral), and their degree of association with each scene entity.
[0100] According to the importance of hot topics, higher initial weights are assigned to the scene entity categories related to them. For example, a higher basic weight value can be assigned to the scene entity categories involved in the hot topics with the highest current popularity. At the same time, the frequency of occurrence of hot topics is taken into consideration. The higher the frequency, the weight of the related scene entity category can be appropriately increased.
[0101] According to the results of sentiment classification, the weights of scene entity categories are dynamically adjusted. For scene entity categories with positive sentiment inclinations, their weights are increased; for scene entity categories with negative sentiment inclinations, their weights are reduced. For example, if a scene entity category has a high positive sentiment score in the barrage, a certain weight value can be increased based on the initial weight.
[0102] Based on the comprehensive calculation of the weights of hot topics and sentiment classification, the final weight of each scenario entity category is obtained. The weighted average method can be adopted. According to the influence degree of hot topics and sentiment classification on the retrieval results, the corresponding weight coefficients are assigned to them, and then the weights of the two are weighted and summed to obtain the comprehensive weight of each scenario entity category. Combine each scenario entity category and its corresponding comprehensive weight to construct a dynamic retrieval weight vector. The dimension of this vector is equal to the number of scenario entity categories, and the value on each dimension represents the weight of the corresponding scenario entity category.
[0103] Comprehensively sort out the preset knowledge graph retrieval strategy and understand various retrieval rules and conditions it contains. For example, the retrieval strategy may stipulate the retrieval methods based on scenario entity categories, attributes, relationships, etc., and the priorities of different retrieval conditions.
[0104] Evaluate the applicability of the preset knowledge graph retrieval strategy in the current live broadcast scenario. Considering the characteristics of the live broadcast (such as real-time and dynamic) and the needs of users (such as obtaining information related to hot topics and sentiment tendencies), judge which retrieval rules and conditions need to be adjusted or retained. According to the dynamic retrieval weight vector, adjust the rules in the preset knowledge graph retrieval strategy. Use the adjusted knowledge graph retrieval strategy to retrieve and filter the scenario entity set. According to the retrieval rules, find the scenario entities that match the current retrieval conditions from the scenario entity set to form the first scenario entity set. Mine the associated historical events of each first scenario entity in the first scenario entity set through the knowledge graph. The knowledge graph stores rich entity relationships and event information, and the historical events related to the entity can be queried in the knowledge graph according to the identifier of the entity.
[0105] Analyze the degree of association between each historical event and the current live hot topic and sentiment tendency. Rules defined manually or methods based on semantic analysis can be used to calculate the degree of association. Utilize the visual attention area detection technology of the blogger to determine the visual attention area of the blogger during the live broadcast. By analyzing information such as the blogger's eye direction and head posture, and combining with the layout of the live broadcast screen, the area of the screen that the blogger is currently focusing on can be roughly determined. Calculate the overlap rate between the bounding box of each first-scene entity and the visual attention area of the blogger. Consider the bounding box of the scene entity and the visual attention area as rectangular regions on a two-dimensional plane, and obtain the overlap rate by comparing the ratio of their overlapping area to their respective areas. Considering comprehensively the associated historical events of each first-scene entity and their degrees of association, as well as the overlap rate between the bounding box and the visual attention area of the blogger, formulate a historical relevance scoring rule. According to the scoring rule, calculate the historical relevance score for each first-scene entity in the first-scene entity set. According to the distribution of the historical relevance scores and business requirements, set an appropriate scoring threshold. Statistical analysis methods can be used to calculate statistics such as the average value and standard deviation of the historical relevance scores, and then determine the threshold based on these statistics. According to the set scoring threshold, screen out the first-scene entities with historical relevance scores higher than the threshold from the first-scene entity set to form a second-scene entity set. These entities are considered to be scene entities highly relevant to the current live hot topic, sentiment tendency, and the blogger's attention area. Integrate and output the second-scene entity set as the dynamic retrieval result.
[0106] In this embodiment, dynamic optimization of knowledge graph retrieval in the tourism live broadcast scenario is achieved, improving the relevance and practicality of the retrieval results.
[0107] In some embodiments, in the above step S104, the IMU unit based on the AR glasses obtains the head posture data of the blogger, and combines the geographic location information to construct a 3D spatial topology map of the tourism live broadcast scenario through the SLAM algorithm, specifically including: The IMU unit based on the AR glasses obtains the head posture data of the blogger, and the head posture data includes three-axis angular velocity and three-axis acceleration; Based on the head posture data, perform quaternion multiplication on the three-axis acceleration using the quaternion differential equation, and combine with the three-axis acceleration for gravity bias compensation to construct a head posture matrix; Obtain the geographic location information of the tourism live broadcast scenario, input the geographic location information and the head posture matrix into the tightly coupled SLAM framework, and achieve multi-source data fusion through the extended Kalman filter to obtain fusion feature data; Based on the fusion feature data, construct a 3D scene topology map using the ORB-SLAM3 framework, and optimize the map structure through the topology consistency detection algorithm.
[0108] In this embodiment, an IMU (Inertial Measurement Unit) with high precision and low latency characteristics is selected and integrated into the AR glasses. The IMU generally includes a three-axis gyroscope and a three-axis accelerometer, which can measure the three-axis angular velocity and the three-axis acceleration respectively.
[0109] According to the real-time requirements of the travel live broadcast scenario, the data acquisition frequency of the IMU unit is set. Generally speaking, a higher acquisition frequency can provide more accurate head pose change information, but it will also increase the data processing burden. The three-axis angular velocity and three-axis acceleration data collected by the IMU unit are transmitted to the data processing module through the communication bus (such as I²C, SPI, etc.) inside the AR glasses.
[0110] Quaternion is a mathematical tool used to represent the rotation in three-dimensional space. Compared with Euler angles and rotation matrices, it has advantages such as high computational efficiency and avoiding gimbal lock. In this embodiment, quaternion is used to describe the head pose change of the blogger.
[0111] According to the three-axis angular velocity data collected by the IMU unit, the value of the quaternion is updated using the quaternion differential equation. The quaternion differential equation describes the relationship between the rate of change of the quaternion over time and the angular velocity. At each sampling moment, according to the current angular velocity value and the previous quaternion state, the quaternion differential equation is solved by numerical integration methods (such as Euler integration, Runge-Kutta integration, etc.) to obtain a new quaternion representation.
[0112] In the stationary state, the three-axis acceleration measured by the IMU unit contains the component of the gravitational acceleration. Since the pose of the AR glasses changes continuously during movement, the components of the gravitational acceleration on each axis will also change accordingly, thus introducing measurement errors. Therefore, it is necessary to compensate for the gravity bias of the three-axis acceleration. Using the head pose information represented by the constructed quaternion, the measured three-axis acceleration is transformed from the sensor coordinate system to the world coordinate system (usually with the direction of gravity as the reference). In the world coordinate system, the direction of the gravitational acceleration is known. By subtracting the components of the gravitational acceleration vector on each axis from the transformed acceleration vector, the acceleration data after removing the gravity bias is obtained. The acceleration data after gravity bias compensation is combined with the head pose information represented by the quaternion to construct a head pose matrix. The head pose matrix is a 3×3 rotation matrix, which describes the rotation relationship between the sensor coordinate system and the world coordinate system. By converting the quaternion into a rotation matrix, the head pose information in three-dimensional space can be obtained.
[0113] Adopt a method of combining multiple positioning technologies to obtain the geographical positioning information of the tourism live broadcast scenario, so as to improve the accuracy and reliability of positioning. For example, simultaneously use GPS (Global Positioning System), Beidou satellite navigation system and base station positioning technology. In an open outdoor environment, mainly rely on GPS or Beidou satellite navigation system to provide high-precision position information; in indoor or areas with severe signal occlusion, use base station positioning technology for assisted positioning.
[0114] Fuse and process the positioning data obtained by different positioning technologies. Since various positioning technologies have different error sources and accuracy characteristics, they are integrated through data fusion algorithms (such as Kalman filter algorithm) to obtain more accurate geographical positioning information.
[0115] Build a tightly coupled SLAM (Simultaneous Localization and Mapping) framework, taking the geographical positioning information and the head pose matrix as input data. The characteristic of the tightly coupled SLAM framework is to closely combine the localization and mapping processes, and improve the accuracy and robustness of the system by simultaneously using multiple sensor data. In the tightly coupled SLAM framework, an Extended Kalman Filter (EKF) is used to achieve multi-source data fusion. The Extended Kalman Filter is a non-linear filtering algorithm that can estimate and update the state of the system. Taking the geographical positioning information and the head pose matrix as observation data, combining the state equation and the observation equation of the system, through the prediction and update steps of the Extended Kalman Filter, continuously optimize the state estimation of the system to obtain the fused feature data. The fused feature data contains more accurate blogger position information and head pose information, as well as environmental feature information extracted from these data.
[0116] ORB-SLAM3 is a feature-point-based SLAM algorithm framework with advantages such as high efficiency and robustness. It realizes the localization and mapping of the environment by extracting feature points (such as ORB feature points) in the environment and tracking the position changes of these feature points in different image frames. Input the fused feature data into the ORB-SLAM3 framework to start constructing a 3D scene topological map. First, extract ORB feature points in each frame of image, and use the head pose matrix and geographical positioning information to determine the positions of these feature points in three-dimensional space. Then, through feature point matching and motion estimation algorithms, calculate the relative motion of the blogger between different positions, and gradually construct a 3D point cloud map of the environment. At the same time, according to the distribution and connection relationship of the feature points, construct the topological structure of the scene, representing the connection relationship between different locations.
[0117] The topological consistency detection algorithm is used to detect and correct possible topological errors in the 3D scene topological map. By analyzing the topological structures of the nodes (representing different locations) and edges (representing the connection relationships between locations) in the map, the algorithm determines whether they conform to the logical relationships of the actual scene. During the construction of the 3D scene topological map, the topological consistency detection algorithm is run regularly. If topological errors are found, such as unreasonable connection relationships between nodes or isolated nodes, the algorithm will correct them according to certain rules. For example, by recalculating the distances and connection relationships between nodes, or deleting incorrect connection edges, the topological structure of the map becomes more accurate and reasonable. After multiple iterations of optimization, a high-quality 3D spatial topological map of the tourism live broadcast scene is obtained.
[0118] In this embodiment, accurate modeling and positioning of the tourism live broadcast scene are achieved, providing richer and more realistic scene information for the tourism live broadcast.
[0119] In some embodiments, in the above step S105, the spatial registration of the dynamic retrieval result and the 3D spatial topological map, combined with the generation of the scene enhancement annotation layer through the NeRF algorithm, specifically includes: Calculating the coordinate correspondence between the dynamic retrieval structure and the 3D spatial topological map through a spatial registration algorithm to construct an affine transformation matrix; Based on the affine transformation matrix, using the NeRF algorithm to calculate the color and density of each coordinate to generate the scene enhancement annotation layer; Dynamically adjusting the annotation style of each annotation box in the scene enhancement annotation layer based on semantic labels.
[0120] In this embodiment, for the dynamic retrieval result, which contains a series of entity information related to the tourism live broadcast scene, these entities have certain position attributes in the virtual space (although they may be represented in a two-dimensional or abstract semantic space). First, this entity information needs to be converted into a form that can be used for feature extraction. For example, if the dynamic retrieval result contains information about a famous scenic spot, the key position points of the scenic spot in the geographical information or virtual scene can be used as candidate feature points. At the same time, auxiliary information such as images and text descriptions related to the scenic spot is collected to more comprehensively describe its features.
[0121] In the 3D spatial topological map, feature point extraction is performed using the geometric and semantic features in the map. Geometric features can be obvious inflection points, corner points, etc. in the map, and semantic features can be the center points or boundary points of different regions (such as buildings, roads, landscapes, etc.) in the map. Through feature extraction algorithms (such as SIFT, SURF, etc., although the formulas are not described, the principle is to use local changes in images or spaces to identify feature points), a set of representative feature points are extracted from the 3D spatial topological map.
[0122] To establish the connection between the dynamic retrieval results and the 3D spatial topological map, the key entities in the dynamic retrieval results are associated with the corresponding feature points in the 3D spatial topological map, and possible matching feature point pairs are searched for between the two.
[0123] Since the dynamic retrieval results and the 3D spatial topological map may use different coordinate systems, it is first necessary to unify their coordinate systems. For example, if the dynamic retrieval results use a two-dimensional coordinate system based on images, while the 3D spatial topological map uses a three-dimensional spatial coordinate system, the two can be made to be in the same coordinate system framework by mapping the two-dimensional coordinate system to a specific plane in the three-dimensional space or projecting the three-dimensional coordinate system onto a two-dimensional plane.
[0124] Select a suitable spatial registration algorithm to calculate the coordinate correspondence between the dynamic retrieval results and the 3D spatial topological map. Commonly used registration algorithms include the Iterative Closest Point (ICP) algorithm, etc. The ICP algorithm continuously iterates to find the best matching relationship between two point sets, minimizing the error between them. During the calculation process, the feature point sets of the dynamic retrieval results and the 3D spatial topological map are used as inputs, and through iterative optimization, the corresponding coordinates of each feature point in the coordinate system of the other are obtained.
[0125] According to the calculated coordinate correspondence, an affine transformation matrix is constructed. The affine transformation matrix is a linear transformation matrix that can describe transformation relationships such as translation, rotation, and scaling between two coordinate systems. By analyzing the position change rules between the feature points, the various parameters in the affine transformation matrix are determined, thus obtaining an affine transformation matrix that can accurately map the dynamic retrieval results to the 3D spatial topological map.
[0126] The NeRF (Neural Radiance Fields) algorithm is a neural network-based scene representation method that can represent a three-dimensional scene as a continuous function, which can output the color and density information of that position according to the coordinates and viewing directions in space. In this embodiment, the NeRF algorithm is used to generate a scene enhancement annotation layer, enabling the annotations to be more realistically integrated into the scene.
[0127] To use the NeRF algorithm, multi-view image data related to the tourism live broadcast scene needs to be collected. These image data can capture the scene from different angles to cover all aspects of the scene. At the same time, the collected image data is preprocessed, including operations such as image correction (such as removing distortion) and color normalization, to improve the data quality. In addition, the image data also needs to be associated with the corresponding spatial coordinates and viewing direction information so that the NeRF algorithm can learn the spatial and color information of the scene.
[0128] Using the previously constructed affine transformation matrix, map the coordinate information in the dynamic retrieval results to the coordinate system of the 3D spatial topological map. Then, sample in the mapped coordinate space and select a series of discrete coordinate points as the input of the NeRF algorithm. These coordinate points cover the distribution area of the dynamic retrieval results in the 3D spatial topological map. Input the sampled coordinate points and the preset viewing direction information into the trained NeRF model. The NeRF model calculates the color and density values of each coordinate point in the given viewing direction according to the scene representation learned inside it. The color value determines the display color of this position during rendering, and the density value affects the degree of light occlusion at this position, thus realizing the three-dimensional representation and rendering of the scene. Generate a scene enhancement annotation layer according to the calculated color and density information of each coordinate point. The annotation layer can be regarded as a transparent layer superimposed on the 3D spatial topological map, where the color and transparency of each pixel point are determined by the color and density of the corresponding coordinate point. In this way, the generated scene enhancement annotation layer can seamlessly blend with the 3D spatial topological map and present a realistic three-dimensional effect according to different viewing angles.
[0129] The dynamic retrieval results usually contain rich semantic information, which can be used as the source of semantic labels. For example, for entities such as scenic spots, buildings, and people in a tourism live broadcast scene, the dynamic retrieval results will give their category, name, feature description and other semantic information. Extract key semantic labels from this information, such as "ancient building", "natural landscape", "tourist", etc. Evaluate the importance of the obtained semantic labels. The importance can be determined according to factors such as the relevance of the semantic label to the current live broadcast theme and user interests, and the frequency of occurrence of the semantic label in the scene. For example, if the current live broadcast theme is about the history and culture of a certain ancient building, then the semantic label "ancient building" has a higher importance; while some semantic labels with a lower frequency of occurrence and less relevance to the theme have relatively lower importance.
[0130] Define multiple different annotation style styles, such as changes in color, shape, size, transparency, etc. For example, for the semantic label "ancient building", it can be stipulated that the color of its annotation box is a simple brown, the shape is a rectangle, and the size is dynamically adjusted according to the importance of the ancient building in the scene; for the semantic label "tourist", the color of the annotation box can be set to a lively yellow, the shape is a circle, and the size is relatively small. Associate the semantic labels with the corresponding annotation style style rules. Dynamically adjust the annotation style style according to the importance evaluation results of the semantic labels. For semantic labels with higher importance, adopt a more prominent and eye-catching annotation style style to attract the user's attention; for semantic labels with lower importance, adopt a relatively low-key and simple annotation style style to avoid causing too much interference to the scene.
[0131] During the process of generating the scene enhancement annotation layer, according to the semantic labels corresponding to each annotation box, the corresponding annotation style rules are applied. By querying the pre-defined rule table, parameters such as color, shape, size, transparency, etc. corresponding to each semantic label are obtained and applied to the drawing of the annotation box. The scene enhancement annotation layer with the applied annotation style is fused and rendered with the 3D spatial topology map to generate the final scene display effect. During the tourism live broadcast, the rendered scene is displayed to the user in real time, enabling the user to intuitively see the tourism live broadcast scene with enhanced annotations and enhancing the user's viewing experience.
[0132] In this embodiment, an enhanced display of the tourism live broadcast scene is realized, improving the information richness of the live broadcast content and the user experience.
[0133] In some embodiments, in the above step S106, the superimposed display of the video stream and the scene enhancement annotation layer and the dynamic adjustment of the annotation visibility using an adaptive transparency algorithm specifically include: Obtain multiple annotation elements of the scene enhancement annotation layer, embed the annotation elements on the frame sequence of the video stream to form target scene data, and the annotation elements include annotation content, bounding box coordinates, and semantic metadata; Obtain the fixation point coordinates of the blogger through an eye movement tracking algorithm, and calculate the fixation focus of each fixation point on the target scene data according to the fixation point coordinates; Calculate the head movement speed based on the head pose data to generate a motion attenuation factor; Determine the annotation content importance by calculating the semantic relevance between the set of hot topics and the semantic metadata; Based on the fixation focus, motion attenuation factor, and annotation content importance, determine the personalized transparency of each annotation element.
[0134] In this embodiment, multiple annotation elements are extracted from the already generated scene enhancement annotation layer. These annotation elements contain rich information. The annotation content clarifies the specific information of the annotated object. For example, in the tourism live broadcast scene, it may annotate "Ancient building name: XX Pavilion", "Scenic spot feature: Historical and cultural relic", etc.; the bounding box coordinates determine the position range of the annotation in the scene, defining the positions of the four corners of the annotation box in the form of coordinate points; the semantic metadata provides deeper semantic information, such as the category of the annotated object (ancient building, natural landscape, etc.), related attributes (construction year, area, etc.).
[0135] Format the extracted annotation elements to facilitate subsequent embedding into the video stream frame sequence. For example, convert the annotation content into text format and determine its basic styles such as font, font size, and color; convert the bounding box coordinates into a coordinate system that matches the video stream frame resolution to ensure that the annotation can be accurately placed at the corresponding position in the video frame.
[0136] Obtain a continuous video stream frame sequence from the video source of the travel live broadcast. The video stream frame sequence is a series of image frames arranged in chronological order, and each frame represents the image information of the live broadcast scene at a certain moment. Adopt a frame-by-frame processing method to embed the formatted annotation elements onto each frame of the video stream frame sequence. For each frame of the image, draw an annotation box at the corresponding position in the image according to the bounding box coordinates of the annotation element, and display the annotation content inside the annotation box. At the same time, associate and store the semantic metadata with the annotation elements for subsequent use. In this way, combine the scene enhancement annotation layer with the video stream frame sequence to form target scene data, so that the video picture contains both the original live broadcast scene and rich annotation information superimposed.
[0137] Select a suitable eye movement tracking device, such as an eye tracker, and install it on the AR glasses. The eye tracker can track the blogger's eye movements in real time and obtain information about the position and gaze direction of the eyes. During the installation process, ensure that the position of the eye tracker is accurate, capable of stably capturing the blogger's eye movement data, and not interfering with the blogger's normal live broadcast operations. During the travel live broadcast, the eye tracker continuously collects the blogger's eye movement data. The eye movement data includes information such as the gaze point coordinates of the eyes, the gaze time, and the blink frequency. Among them, the gaze point coordinates are the key data for calculating the gaze focus, which represents the screen position where the blogger's eyes focus at a certain moment.
[0138] Map the collected gaze point coordinates to the target scene data (i.e., the video frame with annotation elements superimposed). Since the video frame has a specific coordinate system when displayed on the screen, it is necessary to convert the gaze point coordinates collected by the eye tracker into a coordinate system that matches the video frame coordinate system. For each gaze point on the target scene data, analyze its positional relationship with each annotation element. If the gaze point falls within the bounding box of a certain annotation element, it is considered that the annotation element has received the blogger's attention. Calculate the gaze focus of the annotation element under the current gaze point according to factors such as the dwell time and gaze frequency of the gaze point within the annotation element. For example, if an annotation element is gazed at by the blogger for a long time or multiple times, then its gaze focus is relatively high; on the contrary, if it is only briefly gazed at, the gaze focus is relatively low.
[0139] Based on the head posture data collected by the IMU unit of the AR glasses, the movement speed of the head is calculated. The movement speed can be estimated by analyzing the changes in the head posture at different times. For example, based on the changes in the angular velocity and acceleration of the head over a period of time, it can be determined whether the head is stationary, moving slowly, or rotating rapidly. A motion attenuation factor is generated according to the movement speed of the head. When the head moves faster, it means that the blogger may be in a state of quickly browsing the scene. At this time, the visibility of the annotation can be appropriately reduced, so the generated attenuation factor value is smaller; when the head moves slower or is stationary, the blogger may be more focused on the current scene, the visibility of the annotation needs to be improved, and the attenuation factor value is larger. The motion attenuation factor is used to subsequently adjust the transparency of the annotation element to adapt to the different head movement states of the blogger.
[0140] Obtain a set of hot topics from relevant data sources of travel live broadcasts. These hot topics can be hot topics in the current tourism field collected through social media, travel forums, live broadcast platforms and other channels, such as "recommendations of niche attractions in a certain area" and "special food check-in places". Sort and classify the obtained hot topics, remove duplicate or irrelevant topics, and classify the relevant topics for subsequent matching with the semantic metadata of the annotation elements. Match the semantic metadata of each annotation element in the scene enhancement annotation layer with the set of hot topics. Analyze the semantic association between the categories, attributes and other information contained in the semantic metadata and the hot topics. For example, if the semantic metadata of a certain annotation element indicates that it belongs to the category of "niche attractions", and the hot topic set contains the topic of "recommendations of niche attractions in a certain area", then it can be considered that the annotation element has a high semantic relevance with the hot topic. According to the calculation results of the semantic relevance, determine the importance of the annotation content for each annotation element. The higher the semantic relevance of the annotation element, the higher its importance. The importance can be expressed in different levels, such as high, medium, and low, or in a specific numerical range, so as to be used in the subsequent calculation of transparency.
[0141] The calculated gaze focus, motion attenuation factor, and annotation content importance are integrated. These three factors reflect the visibility requirements of annotation elements from the aspects of the blogger's visual attention, head movement status, and the degree of relevance of the annotation content to the hot topic. Different weights are assigned to gaze focus, motion attenuation factor, and annotation content importance according to the actual application scenarios and requirements. For example, if it is believed that the blogger's visual attention has the greatest impact on the visibility of the annotation, a higher weight can be assigned to gaze focus; if the head movement status has a smaller impact on the visibility of the annotation, a lower weight can be assigned to the motion attenuation factor.
[0142] According to the assigned weights, the fixation focus, motion attenuation factor, and annotation content importance are weighted and summed or other comprehensive calculation methods are used to obtain the personalized transparency value of each annotation element. The transparency value is usually between 0 and 1, where 0 represents completely transparent (invisible), and 1 represents completely opaque (fully visible). During the travel live broadcast, the transparency of each annotation element is adjusted in real time according to the calculated personalized transparency value. When the blogger's fixation point falls on a certain annotation element, if the annotation element has a high fixation focus, high importance, and a slow head movement speed (a large motion attenuation factor), its transparency is increased to make it more prominent; conversely, if the fixation focus is low, the importance is low, or the head movement speed is fast (a small motion attenuation factor), its transparency is decreased to reduce the interference to the blogger's vision. In this way, the dynamic adjustment of annotation visibility is achieved, improving the user experience of travel live broadcasts.
[0143] In this embodiment, the intelligent display of annotation information in the travel live broadcast scenario is realized, enabling the blogger to obtain relevant information more conveniently while avoiding the impact of information overload on the live broadcast experience.
[0144] Further, after step S106, the method further includes: Obtain the audience portrait data and real-time feedback data of the live broadcast room; Based on the audience portrait data, determine the interest degree of each user in different contents, and construct an audience attribute vector; Based on the real-time feedback data, determine multiple feedback events to form a feedback event set, where the feedback event includes an event type, an annotation element identifier, and a timestamp; Based on the audience attribute vector and the feedback event set, calculate the display weight of each annotation element through a semantic similarity function based on knowledge graph embedding to obtain a priority score result; Determine the display parameters of each annotation element according to the priority score result to generate a rendering strategy matrix.
[0145] In this embodiment, the audience portrait data is collected from multiple data sources of the live broadcast platform. On the one hand, through the registration information of users on the live broadcast platform, the basic attributes of the audience, such as age, gender, region, etc., are obtained. These information are the basic data for constructing the audience portrait and reflect the basic characteristics of the audience. On the other hand, analyze the historical behavior data of the audience, including the types of live broadcasts watched, the hosts followed, the interactive activities participated in (such as liking, commenting, sharing, etc.), and the duration of staying in the live broadcast room. For example, if a certain audience often watches travel live broadcasts and stays in the travel live broadcast room for a long time, it can be initially judged that the audience has a high interest in travel-related content.
[0146] Utilize the technical architecture of the live streaming platform to capture the feedback data of the audience in real time. When the audience performs operations such as giving likes, commenting, or selecting emojis, the system immediately records the relevant information of these feedback events, including the event type (such as like, comment, emoji selection, etc.), the annotation element identifier (if the feedback is for a certain annotation element in the live streaming room, such as a scenic spot annotation, a product recommendation annotation, etc., then record the unique identifier of this annotation element), and the timestamp (record the specific time when the feedback event occurred).
[0147] According to the business requirements and content characteristics of the live streaming room, divide the live streaming content into multiple categories, such as tourist attraction introductions, food recommendations, cultural and historical explanations, etc. For each content category, define corresponding interest evaluation indicators. For example, for tourist attraction introduction content, the proportion of the viewing duration of this type of live stream by the audience and the frequency of participating in relevant interactions (such as asking questions, commenting on topics related to the scenic spot) can be used as interest evaluation indicators. According to the importance of different content categories to the live streaming room and the degree of attention of the audience, assign weights to each interest evaluation indicator. For example, if tourist attraction introduction is the core content of the live streaming room, then a higher weight can be assigned to the interest evaluation indicators for this content category.
[0148] For each audience, calculate their interest in different content categories based on their historical behavior data and interest evaluation indicators. For example, count the total viewing duration of the audience for tourist attraction introduction live streams in the past period of time and its proportion in the total viewing duration of all live streams, and then combine the number of times the audience participates in relevant interactions to comprehensively calculate the interest score of the audience in tourist attraction introduction content.
[0149] Combine the interest scores of the audience in each content category to construct an audience attribute vector. For example, if there are three content categories: tourist attraction introduction, food recommendation, and cultural and historical explanation, then the audience attribute vector can be expressed as (interest score for tourist attraction introduction, interest score for food recommendation, interest score for cultural and historical explanation). This vector can intuitively reflect the interest preferences of the audience for different content.
[0150] Classify the real-time feedback data collected from the live streaming room and define different feedback event types. Common feedback event types include like events, comment events, emoji selection events, etc. Each event type has its specific meaning and manifestation form. For example, a like event indicates the audience's recognition of the live streaming content, a comment event indicates that the audience has specific opinions or ideas to express, and an emoji selection event can quickly convey the emotional state of the audience. Use natural language processing techniques and pattern recognition algorithms to automatically identify and classify the real-time feedback data. For example, for comment data, the event type it belongs to can be judged by methods such as keyword matching and semantic analysis; for emoji selection data, the event type can be directly determined according to the emoji icon selected by the user.
[0151] For each identified feedback event, record its detailed information, including the event type, annotation element identifier, and timestamp. If the feedback event is for a certain annotation element in the live stream (such as the annotation box of a certain scenic spot), record the unique identifier of this annotation element to facilitate subsequent analysis of the audience's feedback on different annotation elements. The timestamp records the specific time when the feedback event occurred, which helps analyze the timeliness and trend of the feedback event. Integrate all identified feedback events into a feedback event set in chronological order. This feedback event set contains the real-time feedback information of all audiences in the live stream.
[0152] Construct a knowledge graph related to the content of the live stream. First, determine the entities in the knowledge graph, such as the annotation elements in the live stream (scenic spots, products, people, etc.), content categories (travel scenic spot introductions, food recommendations, etc.), and audience attributes (age, gender, interest preferences, etc.). Then, define the relationships between entities. For example, the belonging relationship between an annotation element and a content category (a certain scenic spot belongs to the travel scenic spot introduction category), and the interest association relationship between an audience attribute and a content category (a certain age group of audiences is interested in food recommendation content). Obtain relevant information from the database of the live streaming platform and external data sources (such as travel scenic spot databases, food recommendation databases, etc.) and fill it into the knowledge graph. At the same time, as the live stream content is updated and audience feedback accumulates, regularly update the knowledge graph to ensure that it can accurately reflect the knowledge relationships inside and outside the live stream.
[0153] Using knowledge graph embedding technology, map the entities and relationships in the knowledge graph into a low-dimensional vector space to obtain the embedding representation of each entity. These embedding representations can capture the semantic similarity and relevance between entities. For example, through training the model, entities with similar semantics in the knowledge graph are also closer in the vector space. For each annotation element and audience attribute vector, calculate their semantic similarity in the vector space. Specifically, calculate the similarity between the entity embedding vector corresponding to the annotation element in the knowledge graph and the audience attribute vector. Similarity calculation can use methods such as cosine similarity, and judge their similarity degree by comparing the directions of the two vectors. The higher the similarity, the more matching the annotation element is with the audience's interest preferences.
[0154] Combine the information in the feedback event set to adjust the results obtained from the semantic similarity calculation, and obtain the display weight of each annotation element. For example, if an annotation element is liked or commented on multiple times in the feedback event set, it indicates that the audience has a high degree of attention to this annotation element, and its weight value can be appropriately increased when calculating the display weight. Sort all annotation elements according to the display weight to obtain the priority score result. An annotation element with a high priority score indicates a high match with the audience's interest preferences and a high degree of attention from the audience, and should be preferentially displayed in the live broadcast room.
[0155] According to the display requirements of the live broadcast room and the user experience requirements, determine the display parameters of each annotation element. Common display parameters include display position, display size, display color, display duration, etc. For example, for important annotation elements (such as popular scenic spot annotations), their display position can be set at a prominent position on the screen, the display size can be appropriately enlarged, and the display color can be set to a prominent color to attract the audience's attention.
[0156] According to the priority score result, assign corresponding display parameter values to each annotation element. Annotation elements with a high priority score can obtain more favorable display parameter values, such as a larger display size, a longer display duration, etc.; while annotation elements with a low priority score can appropriately reduce their display size, shorten the display duration, or even not be displayed in some cases.
[0157] Combine all annotation elements and their corresponding display parameter values to construct a rendering strategy matrix. The rendering strategy matrix is a two-dimensional table, where the rows represent different annotation elements and the columns represent different display parameters. Each element in the matrix corresponds to the specific value of an annotation element under a certain display parameter. During the live broadcast, update the rendering strategy matrix in real time according to the audience's feedback and the changes in the live broadcast content. When a new feedback event occurs or the live broadcast content is updated, recalculate the display weight and priority score of the annotation elements, and accordingly adjust the parameter values in the rendering strategy matrix. Apply the updated rendering strategy matrix to the rendering system of the live broadcast room to achieve dynamic display and optimization of the annotation elements.
[0158] In this embodiment, it is possible to intelligently adjust the display method of the annotation elements in the live broadcast room according to the audience's interest preferences and real-time feedback, improving the audience's viewing experience and live broadcast effect.
[0159] Refer to Figure 2 , An embodiment of the present invention provides a blogger live video AR glasses Jingwen annotation system 2, and the system 2 specifically includes: The first annotation module 201 is used to collect the video stream of the travel live broadcast scene in real time through the built-in camera and built-in sensor of the AR glasses, and obtain the live broadcast room interaction data stream in real time through the Apache Flink stream processing engine; The second annotation module 202 is configured to obtain a set of hot topics and a sentiment classification result based on the live streaming interaction data stream by analyzing the bullet screen word frequency, the live streaming time period, and the bullet screen text; The third annotation module 203 is configured to dynamically adjust the knowledge graph retrieval strategy of the travel live streaming scene based on the video stream of the live streaming scene by using the set of hot topics and the sentiment classification result, and obtain a dynamic retrieval result; The fourth annotation module 204 is configured to obtain the blogger's head pose data based on the IMU unit of the AR glasses, and combine the geographic location information to construct a 3D spatial topology map of the travel live streaming scene through the SLAM algorithm; The fifth annotation module 205 is configured to perform spatial registration on the dynamic retrieval result and the 3D spatial topology map, and combine to generate a scene enhancement annotation layer through the NeRF algorithm; The sixth annotation module 206 is configured to superimpose and display the video stream and the scene enhancement annotation layer, and dynamically adjust the annotation visibility by using an adaptive transparency algorithm.
[0160] It can be understood that the content in the embodiment of the blogger live video AR glasses scene text annotation method as Figure 1 shown is applicable to the embodiment of the blogger live video AR glasses scene text annotation system. The functions specifically implemented by the embodiment of the blogger live video AR glasses scene text annotation system are the same as those in the embodiment of the blogger live video AR glasses scene text annotation method as Figure 1 shown, and the beneficial effects achieved are also the same as those in the embodiment of the blogger live video AR glasses scene text annotation method as Figure 1 shown.
[0161] It should be noted that the information interaction, execution process, etc. between the above systems, due to being based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0162] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0163] Referring to Figure 3 , an embodiment of the present invention further provides a computer device 3, including: a memory 302, a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the blogger live video AR glasses Jingwen annotation method as described in any one of the above methods.
[0164] The computer device 3 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that Figure 3 merely an example of the computer device 3, which does not constitute a limitation on the computer device 3, and may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0165] The so-called processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0166] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In some other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 3. Further, the memory 302 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 302 is used to store an operating system, application programs, a Boot Loader, data, and other programs, such as the program code of the computer program. The memory 302 may also be used to temporarily store data that has been output or will be output.
[0167] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the blogger live video AR glasses Jingwen annotation method as described in any one of the above methods.
[0168] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0169] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0170] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0171] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal devices and methods can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.
[0172] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
Claims
1. A method for annotating the scene text of a blogger's live video AR glasses, characterized in that, The method specifically includes: Real-time collect the video stream of the tourism live broadcast scene through the built-in camera and built-in sensors of the AR glasses, and obtain the interactive data stream of the live broadcast room in real time through the Apache Flink stream processing engine; Based on the interactive data stream of the live broadcast room, obtain the hot topic set and sentiment classification result by analyzing the bullet screen word frequency, live broadcast time period, and bullet screen text; Based on the video stream of the live broadcast scene, dynamically adjust the knowledge graph retrieval strategy of the tourism live broadcast scene through the hot topic set and sentiment classification result, and obtain the dynamic retrieval result; Obtain the blogger's head pose data based on the IMU unit of the AR glasses, combine the geographic location information, and construct a 3D spatial topology map of the tourism live broadcast scene through the SLAM algorithm; Perform spatial registration on the dynamic retrieval result and the 3D spatial topology map, and combine to generate a scene enhancement annotation layer through the NeRF algorithm; Overlay and display the video stream and the scene enhancement annotation layer, and dynamically adjust the annotation visibility using the adaptive transparency algorithm.
2. The method according to claim 1, characterized in that, The step of obtaining the hot topic set and sentiment classification result by analyzing the bullet screen word frequency, live broadcast time period, and bullet screen text based on the interactive data stream of the live broadcast room specifically includes: Extract the bullet screen data set from the interactive data stream of the live broadcast room, perform time window segmentation on the bullet screen data set, and divide it into several bullet screen data subsets according to the live broadcast time period; Perform word frequency statistics on the bullet screen data subset within each time window, calculate the heat weight value of each term, and obtain the word frequency vector statistical result; Generate a hot topic set based on the word frequency vector statistical result through an improved LDA topic model; Perform sentiment polarity classification on the bullet screen text of the bullet screen data subset within each time window, and calculate the sentiment score of each bullet screen using the RoBERTa-base model to obtain the sentiment classification result.
3. The method according to claim 1, characterized in that The step of dynamically adjusting the knowledge graph retrieval strategy of the tourism live broadcast scene through the hot topic set and sentiment classification result based on the video stream of the live broadcast scene, and obtaining the dynamic retrieval result specifically includes: Extract the shooting direction angle and timestamp in the video stream of the live broadcast scene, and generate a scene entity set through the object detection model. The scene entity set includes several scene entities and their bounding box coordinates and class labels; Based on the scene entity set, construct a dynamic retrieval weight vector in combination with the hot topic set and sentiment classification result; Based on the dynamic retrieval weight vector, screen out the first scene entity set from the pre-set knowledge graph retrieval strategy; According to the associated historical events and their association degrees of each first scene entity in the first scene entity set, and the overlap rate between the bounding box of each first scene entity and the visual attention area of the blogger, calculate the historical relevance score of each first scene entity in the first scene entity set; Based on the historical relevance score, screen out the second scene entity set from the first scene entity set to obtain the dynamic retrieval result.
4. The method according to claim 1, wherein The step of obtaining the blogger's head pose data based on the IMU unit of the AR glasses, combining the geographic location information, and constructing a 3D spatial topology map of the tourism live broadcast scene through the SLAM algorithm specifically includes: The IMU unit based on the AR glasses acquires the head pose data of the blogger, and the head pose data includes three-axis angular velocity and three-axis acceleration; Based on the head pose data, the three-axis acceleration is subjected to quaternion multiplication operation using the quaternion differential equation, and gravity deviation compensation is combined with the three-axis acceleration to construct a head pose matrix; Obtain the geolocation information of the travel live broadcast scene, input the geolocation information and the head pose matrix into the tightly coupled SLAM framework, and realize multi-source data fusion through the extended Kalman filter to obtain fused feature data; Based on the fused feature data, use the ORB-SLAM3 framework to construct a 3D scene topology map, and optimize the map structure through the topology consistency detection algorithm.
5. The method according to claim 1, characterized in that The spatial registration of the dynamic retrieval result and the 3D spatial topology map, combined with the generation of the scene enhancement annotation layer through the NeRF algorithm, specifically includes: Calculate the coordinate correspondence between the dynamic retrieval structure and the 3D spatial topology map through the spatial registration algorithm to construct an affine transformation matrix; Based on the affine transformation matrix, use the NeRF algorithm to calculate the color and density of each coordinate to generate a scene enhancement annotation layer; Dynamically adjust the annotation style of each annotation box in the scene enhancement annotation layer based on semantic tags.
6. The method according to claim 1, wherein The overlay display of the video stream and the scene enhancement annotation layer, and the dynamic adjustment of the annotation visibility using the adaptive transparency algorithm, specifically includes: Obtain multiple annotation elements of the scene enhancement annotation layer, embed the annotation elements on the frame sequence of the video stream to form target scene data, and the annotation elements include annotation content, bounding box coordinates, and semantic metadata; Obtain the fixation point coordinates of the blogger through the eye movement tracking algorithm, and calculate the fixation focus of each fixation point on the target scene data according to the fixation point coordinates; Calculate the head movement speed based on the head pose data to generate a motion attenuation factor; Determine the importance of the annotation content by calculating the semantic correlation between the set of hot topics and the semantic metadata; Based on the fixation focus, motion attenuation factor, and importance of the annotation content, determine the personalized transparency of each annotation element.
7. The method according to claim 6, characterized in that, The method further includes: Obtain the audience portrait data and real-time feedback data of the live broadcast room; Based on the audience portrait data, determine the interest degree of each user in different contents, and construct an audience attribute vector; Based on the real-time feedback data, determine multiple feedback events to form a feedback event set, and the feedback events include event type, annotation element identifier, and timestamp; Based on the audience attribute vector and the feedback event set, calculate the display weight of each annotation element through the semantic similarity function based on knowledge graph embedding to obtain a priority score result; Determine the display parameters of each annotation element according to the priority score result to generate a rendering strategy matrix.
8. A blogger live video AR glasses scene text annotation system, characterized in that, The system specifically includes: The first annotation module is used to collect the video stream of the travel live broadcast scene in real time through the built-in camera and built-in sensor of the AR glasses, and obtain the live broadcast room interaction data stream in real time through the Apache Flink stream processing engine; The second annotation module is used to obtain a set of hot topics and sentiment classification results based on the live broadcast room interaction data stream by analyzing the bullet screen word frequency, live broadcast time period, and bullet screen text; The third annotation module is used to dynamically adjust the knowledge graph retrieval strategy of the tourism live broadcast scenario based on the video stream of the live broadcast scenario, and obtain dynamic retrieval results through the hot topic set and the sentiment classification result; The fourth annotation module is used to obtain the blogger's head pose data based on the IMU unit of the AR glasses, combine the geographic location information, and construct a 3D spatial topology map of the tourism live broadcast scenario through the SLAM algorithm; The fifth annotation module is used to perform spatial registration on the dynamic retrieval result and the 3D spatial topology map, and combine to generate a scene enhancement annotation layer through the NeRF algorithm; The sixth annotation module is used to superimpose and display the video stream and the scene enhancement annotation layer, and dynamically adjust the annotation visibility by using the adaptive transparency algorithm.
9. A computer device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, it implements the blogger live video AR glasses scene text annotation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon. When the computer program is run by the processor, it implements the blogger live video AR glasses scene text annotation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Live broadcast data processing method and device, electronic equipment and storage medium
CN111343467A
Live broadcast method and system based on artificial intelligence
CN118450156A
Interaction system applied to multi-scene digital human
CN119052521A
Automatic labeling and acquisition standardization method and system based on video content
CN119583881A
Cited By
Accompanying dialogue system and method based on real-time environment perception and knowledge graph enhancement
CN120894187A
Method and system for monitoring and processing linkage information of multiple live broadcasting rooms
CN121547606A