Method, device and system for multi-modal interactive perception of dynamic panoramic environmental information
By employing a multimodal interactive perception method, audio and tactile feedback are used to provide visually impaired users with personalized environmental exploration, understanding, and social interaction. This solves the problem that traditional assistive tools cannot provide in-depth interaction, thereby improving the independence and quality of life of visually impaired users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-04-18
- Publication Date
- 2026-08-04
AI Technical Summary
Traditional assistive tools such as canes and guide dogs cannot provide in-depth, interactive environmental exploration, causing blind and low-vision people to passively receive path descriptions in unfamiliar environments, unable to actively explore, understand, and socially interact, resulting in feelings of helplessness and excessive cognitive load.
This paper provides a multimodal interactive perception method for dynamic panoramic environmental information. It recommends environmental objects to visually impaired users through multimodal interaction (audio guidance and tactile feedback), generates a personalized recommendation list based on dynamic panoramic video using an environmental recommendation model, and iteratively updates the list based on user feedback. It supports exploration, understanding, recall, and social interaction.
Enhance the independence and quality of life of visually impaired users by providing personalized environmental perception assistance to improve their ability to actively explore, understand, and interact with others.
Smart Images

Figure CN120491806B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multimodal interactive perception method, device and system for dynamic panoramic environmental information. Background Technology
[0002] With the continuous advancement of technology, more and more technologies are being applied to improve the quality of life for people with disabilities. For blind and low-vision individuals (BLV), exploring and perceiving unfamiliar environments has always been a significant challenge. Traditional assistive tools, such as canes and guide dogs, while helping BLV individuals with daily mobility to some extent, are clearly insufficient in providing environmental information and enhancing the experience. Blind people can generally only passively receive basic path descriptions and are unable to engage in in-depth, interactive environmental exploration. Summary of the Invention
[0003] To address the technical problems existing in the prior art, this invention provides a multimodal interactive perception method, device, and system for dynamic panoramic environmental information, which can dynamically adapt to user preferences and provide personalized environmental perception assistance, thereby enhancing the independence and quality of life of visually impaired users.
[0004] This invention provides a multimodal interactive perception method for dynamic panoramic environmental information, applied to a user terminal device. The method includes: acquiring a recommendation list of environmental objects for perception; recommending environmental objects to a visually impaired user through multimodal interaction based on the recommendation list, thereby assisting the visually impaired user in environmental perception from multiple dimensions; the multimodal interaction includes audio guidance and tactile feedback; the multiple dimensions include exploration, understanding, recall, and social interaction dimensions; wherein, the recommendation list is a list generated by an environmental recommendation model based on a dynamic panoramic environmental video, sorting several environmental objects according to a comprehensive interest score; the comprehensive interest score is a score that comprehensively considers the aesthetic, novelty, and needs factors of the visually impaired user; the environmental recommendation model is used to iteratively update the association weights of aesthetic, novelty, and needs factors based on the tactile feedback, thereby iteratively updating the comprehensive interest score and the recommendation list.
[0005] According to a multimodal interactive perception method for dynamic panoramic environmental information provided by the present invention, when a visually impaired user is exploring the environment, information about recommended environmental objects is described to the visually impaired user via audio based on a recommendation list; the information about the recommended environmental objects includes the name, location, and attributes of the recommended environmental objects; the exploration actions of the visually impaired user are determined; the exploration actions are used to characterize the visually impaired user's preference information for the currently recommended environmental objects; the preference information includes liking and disliking; and the recommendation list is dynamically adjusted based on the exploration actions.
[0006] According to a multimodal interactive perception method for dynamic panoramic environment information provided by the present invention, when a visually impaired user is understanding the environment, based on the recommendation list, a two-layer hierarchical scene interaction architecture is adopted to form several environmental objects in a main layer graph structure and a sub-layer graph structure. The nodes of the main layer graph structure are coarse-grained environmental objects, and the edges of the main layer graph structure represent the relationships between the coarse-grained objects. The nodes of the sub-layer graph structure are fine-grained environmental objects, and the edges of the sub-layer graph structure represent the relationships between the fine-grained environmental objects. Based on the recommendation list, the information of the coarse-grained environmental objects in the main layer graph structure is described to the visually impaired user via audio. The understanding action of the visually impaired user is determined. The understanding action is used to characterize the visually impaired user's need to continue to understand the information of finer-grained environmental objects under the currently recommended coarse-grained environmental objects. Based on the understanding action, the information of the fine-grained environmental objects in the sub-layer graph structure corresponding to the currently recommended coarse-grained environmental objects is described to the visually impaired user via audio.
[0007] According to a multimodal interactive perception method for dynamic panoramic environmental information provided by the present invention, when a visually impaired user recalls the environment, the method determines the recall action of the visually impaired user; the recall action is used to characterize the visually impaired user's need to recall historical environmental objects; based on the recall action, historical multimodal interactive environmental perception information is retrieved, and the information of the historical environmental objects is described to the visually impaired user via audio.
[0008] According to a multimodal interactive perception method for dynamic panoramic environmental information provided by the present invention, when a visually impaired user engages in environmental social interaction, the method determines the social interaction action of the visually impaired user; the social interaction action represents the current visually impaired user's need to share information about its historical environmental objects with another visually impaired user; based on the social interaction action, the method retrieves historical multimodal interactive environmental perception information and sends the historical multimodal interactive environmental perception information to the receiving device of the other visually impaired user, so that the current visually impaired user and the other visually impaired user can share the information about the historical environmental objects.
[0009] According to the present invention, a multimodal interactive perception method for dynamic panoramic environment information further includes: converting the dynamic panoramic environment video into a semantic graph sequence using a scene graph generation algorithm, and inputting the semantic graph sequence into the environment recommendation model; the environment recommendation model is trained based on environment video training samples using graph mask self-supervised learning and a multimodal attention mechanism; the environment recommendation model includes a background network, an aesthetic network, a novelty network, and a demand network; the background network is used to identify background objects in the semantic graph sequence through a background attention mechanism and calculate the score of the background object to obtain a background score; the aesthetic network is used to identify... The semantic graph sequence continuously identifies objects of interest that appear within a preset time period, and calculates a score for each object of interest to obtain an aesthetic score. The novelty network identifies objects of interest that appear for the first time compared to previous frames in the semantic graph sequence using a novelty attention mechanism, and calculates a score for each object of interest to obtain a novelty score. The demand network identifies objects of demand related to the physiological and safety needs of the visually impaired user in the semantic graph sequence using a demand attention mechanism, and calculates a score for each object of demand to obtain a demand score. The comprehensive interest score is a score determined by combining the background score, the aesthetic score, the novelty score, and the demand score.
[0010] According to a multimodal interactive perception method for dynamic panoramic environment information provided by the present invention, the environment recommendation model further includes a user interaction adapter; the user interaction adapter is used to dynamically adjust the weights of the background score, the aesthetic score, the novelty score, and the demand score based on the tactile feedback using a maximum likelihood estimation algorithm.
[0011] This invention also provides a multimodal interactive sensing device for dynamic panoramic environment information, applied to a user terminal device. The device includes: a recommendation list acquisition module for acquiring a recommendation list of environmental objects for perception; and an environmental object recommendation module for recommending environmental objects to visually impaired users through multimodal interaction based on the recommendation list, thereby assisting the visually impaired users in environmental perception from multiple dimensions. The multimodal interaction includes audio guidance and tactile hierarchical feedback; the multiple dimensions include exploration, understanding, recall, and social interaction dimensions. The recommendation list is a list generated by an environmental recommendation model based on a dynamic panoramic environment video, sorting several environmental objects according to a comprehensive interest score. The comprehensive interest score is a score that comprehensively considers the aesthetic, novelty, and needs factors of the visually impaired user. The environmental recommendation model is used to iteratively update the association weights of aesthetic, novelty, and needs factors based on the tactile hierarchical feedback, thereby iteratively updating the comprehensive interest score and the recommendation list.
[0012] This invention also provides a multimodal interactive perception system for dynamic panoramic environmental information, comprising: a data acquisition device for capturing dynamic panoramic environmental video and transmitting the dynamic panoramic environmental video to a server; the server for generating a recommendation list based on an environment recommendation model running on the dynamic panoramic environmental video; the recommendation list being a list of several environmental objects sorted according to a comprehensive interest score; the comprehensive interest score being a score that comprehensively considers aesthetics, novelty, and needs of visually impaired users; the environment recommendation model being used to iteratively update the association weights of aesthetics, novelty, and needs based on the tactile feedback hierarchy, so as to iteratively update the comprehensive interest score and the recommendation list; and a user terminal device for executing the above-described multimodal interactive perception method for dynamic panoramic environmental information.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a multimodal interactive perception method for dynamic panoramic environment information as described above.
[0014] This invention provides a multimodal interactive perception method, device, and system for dynamic panoramic environmental information. The method includes acquiring a recommended list of environmental objects, and then recommending these objects to visually impaired users through multimodal interaction based on the list. The multimodal interaction includes audio guidance and tactile feedback, covering multiple dimensions such as exploration, understanding, recall, and social interaction. The recommended list is generated by an environmental recommendation model based on dynamic panoramic video. This model sorts multiple environmental objects according to a comprehensive interest score. The comprehensive interest score considers the visually impaired user's ratings of aesthetics, novelty, and needs. The environmental recommendation model uses tactile feedback to iteratively update the association weights of aesthetics, novelty, and needs, thereby iteratively updating the comprehensive interest score and the recommended list. This invention can dynamically adapt to user preferences, providing personalized environmental perception assistance, thereby enhancing the independence and quality of life of visually impaired users. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a multimodal interactive perception method for dynamic panoramic environmental information provided by the present invention.
[0017] Figure 2This is a schematic diagram of the structure of a multimodal interactive sensing device for dynamic panoramic environmental information provided by the present invention.
[0018] Figure 3 This is a schematic diagram of the structure of a multimodal interactive sensing system for dynamic panoramic environmental information provided by the present invention.
[0019] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] Contact with the natural environment is crucial for an individual's physical and mental well-being. However, millions of blind and visually impaired (BLV) individuals worldwide also yearn to actively explore these unknown wonders. Yet, existing research often assumes that BLV individuals have clearly defined needs, focusing solely on functional assistance such as navigation and obstacle avoidance (e.g., guide dogs). This leads to their passive dependence on others in unfamiliar environments, resulting in strong feelings of helplessness and severely hindering their independent exploration during dynamic sightseeing, post-experience memory retention, and interaction with other BLV individuals. The vast amount of visual information in unfamiliar environments far exceeds the perceptual capacity of BLV individuals, leading to excessive cognitive load. Therefore, helping BLV individuals understand and enjoy unfamiliar environments is an urgent need.
[0022] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a multimodal interactive perception method for dynamic panoramic environmental information provided by the present invention.
[0023] This invention provides a multimodal interactive sensing method for dynamic panoramic environmental information, applied to a user terminal device 3, the method comprising: 101: Obtain a recommended list of objects in the environment; 102: Based on a recommendation list, environmental objects are recommended to visually impaired users through multimodal interaction to assist them in environmental perception from multiple dimensions. Multimodal interaction includes audio guidance and tactile feedback. The multiple dimensions include exploration, understanding, recall, and social interaction. The recommendation list is a list generated by the environmental recommendation model based on dynamic panoramic environmental videos, which sorts several environmental objects according to a comprehensive interest score. The comprehensive interest score is a score that combines the aesthetic, novelty, and demand factors of visually impaired users. The environmental recommendation model is used to iteratively update the association weights of aesthetic, novelty, and demand factors based on tactile feedback, thereby iteratively updating the comprehensive interest score and the recommendation list.
[0024] This embodiment details the implementation of a multimodal interactive perception method for dynamic panoramic environmental information. This method aims to provide visually impaired users with a platform to actively explore and deeply understand unfamiliar environments, achieving long-term memory and effective communication. It reconstructs massive amounts of information in complex scenes into structured, personalized semantic graphs, enhancing their independence and quality of life.
[0025] This invention integrates a lightweight, portable consumer-grade device combination. First, a wearable acquisition device 1, such as a panoramic camera, captures real-time panoramic video of the user's surrounding environment. This video data is transmitted wirelessly to a server 2. The server 2 can be a laptop computer connected to a local area network (LAN), ensuring high communication reliability with virtually no latency or interference. An environmental recommendation model deployed on server 2 utilizes image recognition and artificial intelligence technologies to analyze environmental objects in the video stream, such as trees, buildings, and pedestrians, ranking multiple environmental objects according to a comprehensive interest score. This comprehensive interest score considers the visually impaired user's aesthetics, novelty, and needs to provide a personalized recommendation list. The recommendation list is then sent to the user's terminal device, such as a smartphone or a dedicated assistive device. Based on the recommendation list, the user's terminal device 3 recommends environmental objects to the user through a multimodal interaction paradigm of audio guidance and haptic feedback, covering four dimensions: exploration, understanding, recall, and social interaction. Audio guidance provides the user with real-time environmental information, such as object names, locations, and features. For example, when a visually impaired user is walking in an outdoor environment, an audio description is provided through headphones, such as "There is a large tree 10 meters ahead." Haptic feedback provides real-time information about the user's interaction with the environment through vibration or touch responses on the user terminal device 3. Users provide feedback on recommended objects through interaction with the user terminal device 3, such as by touching the screen or issuing voice commands. This feedback is used to update the environmental recommendation model on server 2. The environmental recommendation model iteratively updates the association weights of aesthetics, novelty, and demand factors based on the user's haptic feedback, further optimizing the comprehensive interest score and recommendation list. This means that as the user interacts with the user terminal device 3, the model can learn the user's preferences and adjust the recommended environmental objects accordingly, making the recommendations more personalized and accurate. For example, assuming a visually impaired user is exploring a park, the user terminal device 3 might recommend "the bench ahead" and "the fountain to the left." The user can obtain more information, such as the bench's material and design details, by touching the bench icon. Simultaneously, the user terminal device 3 will adjust the weights of future recommendations based on the user's interaction feedback, placing objects that the user is more interested in higher in future recommendations.
[0026] Furthermore, the environmental recommendation model can learn user behavior patterns in different environments and automatically adjust recommendation strategies to adapt to different scenarios. This adaptive learning capability allows the environmental recommendation model to continuously optimize over time, providing users with a more personalized and accurate environmental perception experience. In this way, visually impaired users can navigate and explore their environment with greater confidence and independence.
[0027] It's important to note that modern smartphones possess powerful computing capabilities, theoretically allowing environment recommendation models to run entirely on smartphones without the need for laptops. However, this requires significant engineering work to optimize mobile deployment, which is a crucial direction for future efforts to make user terminal devices more practical and accessible for real-world applications.
[0028] Experiments show that the method of this invention, by guiding visually impaired users through an exploration path and dynamically adjusting information supply, enhances their positive emotions and memory accuracy. Both emotional valence (pleasure) and arousal (activation) are significantly improved, bringing unprecedented pleasure and unforgettable experiences to visually impaired users. This invention opens the door for visually impaired individuals to independently explore the world, propelling the construction of an accessible society into a new stage.
[0029] In a preferred embodiment, when a visually impaired user is exploring their environment, information about recommended environmental objects is described to the visually impaired user via audio based on a recommendation list. The information about the recommended environmental objects includes their names, locations, and attributes. The exploration actions of the visually impaired user are determined. These exploration actions are used to characterize the visually impaired user's preference information for the currently recommended environmental objects. The preference information includes "likes" and "dislikes." The recommendation list is dynamically adjusted based on the exploration actions.
[0030] In this embodiment, when a visually impaired user explores the environment, server 2 first generates a recommendation list using an environmental recommendation model. This list, based on dynamic panoramic environmental video, sorts multiple environmental objects according to a comprehensive interest score. The comprehensive interest score integrates the visually impaired user's aesthetic, novelty, and needs factors, significantly enhancing the immersive experience. Then, user terminal device 3 describes the recommended environmental objects to the visually impaired user via audio, including the object's name, location, and attributes. For example, user terminal device 3 might play the audio: "There is a bench to your left, about two meters away, made of wood." Simultaneously, user terminal device 3 monitors the visually impaired user's exploration actions to determine their preference information for the currently recommended environmental objects. Exploration actions may include, but are not limited to, touching, clicking, or performing specific gestures on the device; these actions represent the user's liking or disliking of a particular environmental object. For example, if the user actively touches or clicks the bench icon after hearing the description, user terminal device 3 records this action as preference information.
[0031] The exploration interface of the user terminal device 3 in this embodiment is designed specifically for visually impaired users. It provides real-time prompts for interesting, novel, or necessary environmental objects in the surrounding environment, allowing users to receive tactile feedback based on their preferences for the currently recommended environmental objects, thereby enhancing the user's interactive experience with the environment. Employing tactile feedback technology, it allows users to perceive and select environmental objects by touching the screen. The interface design takes into account the special needs of visually impaired users, using a simple and intuitive layout, and distinguishing different environmental objects through different tactile patterns or vibration modes. Users browse the list of recommended environmental objects on the exploration interface by touching and swiping. Each object is associated with a specific tactile feedback mode; for example, trees may be associated with a rough touch, while water may be associated with a smooth, undulating touch. Users can obtain detailed descriptions of these objects by touching them, such as the object's name, location, and characteristics.
[0032] Based on these exploratory actions, the recommendation list can be dynamically adjusted. If a user shows a liking action towards an environmental object, the object's weight in the recommendation list will increase, making it rank higher in future recommendations. Conversely, if a user shows a dislike action towards an object, its weight will decrease, reducing its frequency of appearance in future recommendations. This dynamic adjustment mechanism can better adapt to users' individual preferences, providing a more personalized environmental awareness experience. For example, a user might be particularly interested in a newly discovered coffee shop (novelty) or have an urgent need for a nearby restroom (need). Users can express their preferences through simple tactile actions (such as long press or double tap). This preference feedback is sent back to server 2 in real time to update the environmental recommendation model.
[0033] Furthermore, it can learn users' general preferences for different types of environmental objects and adjust the recommendation algorithm accordingly. For example, if user terminal device 3 discovers that users typically show high interest in trees and bodies of water, it may give these types of objects higher initial weights when generating the recommendation list.
[0034] As a preferred embodiment, when visually impaired users are understanding their environment, a two-layer hierarchical scene interaction architecture is adopted based on a recommendation list to form several environmental objects in a main-layer graph structure and a sub-layer graph structure. Nodes in the main-layer graph structure represent coarse-grained environmental objects, and edges represent the relationships between these coarse-grained objects. Nodes in the sub-layer graph structure represent fine-grained environmental objects, and edges represent the relationships between these fine-grained objects. Based on the recommendation list, information about the coarse-grained environmental objects in the main-layer graph structure is described to the visually impaired user via audio. The user's understanding actions are determined. These understanding actions represent the user's need to further understand information about finer-grained environmental objects under the currently recommended coarse-grained environmental objects. Based on these understanding actions, information about the fine-grained environmental objects in the sub-layer graph structure corresponding to the currently recommended coarse-grained environmental object is described to the visually impaired user via audio.
[0035] In this embodiment, user terminal device 3 first adopts a two-layer hierarchical scene interaction architecture based on a recommendation list to improve visually impaired users' cognition and understanding of environmental information. The hierarchical graph structure promotes semantic association memory and supports post-event social sharing. In the main layer graph structure, nodes represent coarse-grained environmental objects, such as trees and buildings, which are the main elements that users first encounter during environmental exploration. The edges of the main layer graph structure represent the relationships between these coarse-grained objects, such as proximity or containment relationships. The sub-layer graph structure contains fine-grained environmental objects, which provide more detailed information, such as the texture of leaves and the windows of buildings. The edges of the sub-layer graph structure represent the relationships between these fine-grained environmental objects, such as the relationship between parts and the whole.
[0036] User terminal device 3 describes information about coarse-grained environmental objects in the main layer diagram structure to the visually impaired user via audio, such as: "There is a big tree in front of you with a thick trunk and lush leaves." Users express their interest in finer-grained information by touching or clicking nodes in the main layer diagram, and these actions are recognized as understanding actions.
[0037] Based on these understanding actions, user terminal device 3 uses audio to describe to the visually impaired user the information of fine-grained environmental objects in the sub-layer structure corresponding to the currently recommended coarse-grained environmental object, such as: "The leaves of the big tree are green, shaped like a palm, and may feel a little rough." Such descriptions help users build a deeper understanding of the environment.
[0038] The user terminal device 3 also includes a graph structure recognition enhancement interface in its interface design, which extracts complex environmental information into a sparse hierarchical graph structure. The graph structure recognition enhancement interface includes a main-layer graph structure interface and a sub-layer graph structure interface. The main-layer graph structure interface is a simplified information presentation layer interface formed by mapping the recommendation list to the scene topology. In this interface, environmental objects are organized into a simplified graph structure. This simplified graph structure helps visually impaired users quickly grasp the overall layout and key objects of the environment. The sub-layer graph structure interface is an interface that displays specific nodes and their relationships after visually impaired users autonomously choose to zoom in on specific nodes in the main-layer graph structure interface according to their perceived scene hierarchy. When a user is interested in a node in the main-layer graph structure, they can zoom in on that node through touch or voice commands. The sub-layer graph structure interface then displays detailed information about that node and its relationships with other nodes, such as adjacent paths and nearby facilities. This two-layer hierarchical scene interaction architecture allows visually impaired users to autonomously choose the display level and level of detail of information according to their needs and interests. Users can start with the macroscopic main layer diagram structure interface and gradually delve into the microscopic sub-layer diagram structure interface to gain a more detailed understanding of specific environmental objects and their associated information.
[0039] In a preferred embodiment, when a visually impaired user is recalling their environment, the user's recall actions are determined; the recall actions are used to characterize the visually impaired user's need to recall historical environmental objects; based on the recall actions, historical multimodal interactive environmental perception information is retrieved, and information about historical environmental objects is described to the visually impaired user via audio.
[0040] In this embodiment, the user terminal device 3 is equipped with advanced data storage and retrieval functions, capable of recording every user interaction, including touch operations, audio feedback selection, and haptic feedback preference settings, supporting the user's immersive experience during the guided tour. This data is correlated with the user's environmental perception information at specific times and locations, forming a historical record of multimodal interactive environmental perception information.
[0041] User terminal device 3 first needs to recognize the visually impaired user's recall actions. These actions may include touching a specific area on the screen, issuing a specific voice command, or performing a preset gesture to indicate that the user wants to recall an environmental object encountered during a previous exploration. For example, the user may trigger a recall action by double-tapping the screen or saying "I want to recall that fountain."
[0042] Once the user terminal device 3 recognizes the recall action, it retrieves relevant multimodal interactive environmental awareness information from stored historical data. This information may include audio descriptions of environmental objects, haptic feedback, and the user's interaction history with these objects. The user terminal device 3 integrates this information and prepares to describe the historical environmental objects to the user via audio, allowing the user to relive the environmental awareness of that time.
[0043] For example, if a user recalls pointing to a fountain in a park they have previously explored, the user terminal device 3 will retrieve audio descriptions and haptic feedback related to the fountain from historical data, and then describe the fountain's features to the user through an audio player, such as "The fountain you encountered before is located in the center of the park, with water jets about three meters high, and is surrounded by smooth stone benches."
[0044] In a preferred embodiment, when a visually impaired user engages in environmental social interaction, the social interaction action of the visually impaired user is determined; the social interaction action is used to represent the current visually impaired user's need to share information about their historical environmental objects with another visually impaired user; based on the social interaction action, historical multimodal interactive environmental perception information is retrieved and sent to the receiving device of another visually impaired user, so that the current visually impaired user and the other visually impaired user can share information about historical environmental objects.
[0045] Visually impaired individuals have a strong desire to share experiences and a sense of community. In this embodiment, the user terminal device 3 first needs to recognize the social interaction actions of visually impaired users. These actions may include specific touch gestures, voice commands, or device vibrations, indicating that the user wishes to share their historical environmental object information with another visually impaired user. For example, a user may trigger a social interaction action by performing a specific gesture on the touchscreen or saying "share this scene."
[0046] Once user terminal device 3 recognizes a social interaction action, it automatically retrieves relevant multimodal interactive environmental awareness information from historical data. This information includes audio descriptions of environmental objects, haptic feedback, and the user's interaction history with these objects. User terminal device 3 packages this information into a shared packet and transmits it wirelessly to the receiving device of another visually impaired user, such as a smartphone, tablet, or other wearable device.
[0047] After receiving information, the receiving device will play an audio description of the environmental objects through its audio output device. It may also provide additional sensory information through haptic feedback devices, allowing another visually impaired user to experience the shared environmental objects. Furthermore, the receiving device can offer interactive feedback options, allowing users to request more information or comment on the shared content, thereby promoting social interaction between the two visually impaired users. It supports users in transforming fragments of their journeys and emotional memories into shareable digital content. This function not only promotes knowledge exchange within the visually impaired community but also strengthens social connections through collective memory, forming a unique cultural interaction space.
[0048] For example, if a visually impaired user encounters an interesting sculpture while exploring a park and wants to share their discovery with a friend, they can perform a social interaction action. The user's terminal device 3 will recognize this action, retrieve detailed information about the sculpture from previous explorations, and then send this information to the friend's device. The friend can then hear a description of the sculpture and may even feel its surface texture through tactile feedback, thus experiencing this environmental object without being physically present.
[0049] This memory-based retrieval function not only helps visually impaired users better recall and understand their experiences, but also enhances their social interactions with others. By reviewing and sharing past environmental perception information, visually impaired users can participate more deeply in social activities, improving their social engagement and quality of life. Furthermore, this function can also be used for educational and training purposes, helping visually impaired users learn how to better utilize user terminal devices 3 to explore and understand their environment.
[0050] As a preferred embodiment, the method further includes: converting the dynamic panoramic environment video into a semantic graph sequence using a scene graph generation algorithm, and inputting the semantic graph sequence into an environment recommendation model; the environment recommendation model is trained based on environment video training samples using graph mask self-supervised learning and a multimodal attention mechanism; the environment recommendation model includes a background network, an aesthetic network, a novelty network, and a demand network; the background network is used to identify background objects in the semantic graph sequence through a background attention mechanism and calculate the score of the background objects to obtain a background score; the aesthetic network is used to identify objects of interest that continuously appear within a preset time period in the semantic graph sequence through an aesthetic attention mechanism and calculate the score of the objects of interest to obtain an aesthetic score; the novelty network is used to identify objects of novelty that appear for the first time compared to previous frames in the semantic graph sequence through a novelty attention mechanism and calculate the score of the novelty objects to obtain a novelty score; the demand network is used to identify objects of demand related to the physiological and safety needs of visually impaired users in the semantic graph sequence through a demand attention mechanism and calculate the score of the objects of demand to obtain a demand score; the comprehensive interest score is a score determined by combining the background score, aesthetic score, novelty score, and demand score.
[0051] In this embodiment, dynamic panoramic environment video samples are converted into semantic graph sequences using a scene graph generation algorithm. The scene graph generation algorithm can identify key objects in the video and construct a graph structure describing these objects and their relationships. These spatiotemporal semantic graph sequences from first-person panoramic videos are used as input data to train an environment recommendation model. To train the environment recommendation model, a graph masking self-supervised learning method is employed. In this process, a subset of graph nodes (i.e., environment objects) are randomly selected for masking, temporarily hiding their information. The model's goal is to predict the features of masked nodes based on the information of the unmasked nodes. This method enables the model to learn the intrinsic connections between nodes and the semantic structure of the environment, effectively eliminating the subjective aesthetic bias of manual annotation. The model output generates key environment objects that dynamically highlight the scene.
[0052] In unfamiliar environments, it is crucial to provide effective information while avoiding cognitive overload. Visually impaired users highly value aesthetics, novelty, and needs when exploring their environment, but aesthetic judgments exhibit significant individual differences. For example, some may prefer magnificent scenery, while others appreciate history and culture, and still others enjoy the sounds of cicadas and birds. To address this, the environment recommendation model is based on environmental video training samples, employing a multimodal attention mechanism to refine information and a graph mask self-supervised learning mechanism for training, effectively eliminating subjective biases from manual annotation. These environmental video training samples can be tens of thousands of publicly available travel videos, allowing the model to capture semantic symbiotic relationships among thousands of environmental objects. The training of the environment recommendation model reveals unique latent contextual information during sighted individuals' travel experiences. For instance, in a park visit scenario, the embedding of "duck" is closer to water-related categories such as "dock," "pond," "fish," and "bridge," demonstrating the contextual association between "duck" and "pond" in a specific scenario.
[0053] The environmental recommendation model consists of four sub-networks: background network, aesthetic network, novelty network, and demand network.
[0054] The background network identifies background objects, such as the sky and trees, in a semantic graph sequence through a background attention mechanism, and calculates scores for these background objects to obtain a background score. This score reflects the stability and prevalence of the objects as background.
[0055] The aesthetic network identifies objects of interest, such as sculptures and fountains, that continuously appear in a semantic graph sequence within a preset time period through an aesthetic attention mechanism, and calculates a score for these objects of interest, resulting in an aesthetic score. This score reflects the aesthetic and interest value of the object to the user.
[0056] The Novelty Network identifies novel objects that appear for the first time in a semantic graph sequence compared to previous frames through a novelty attention mechanism, and calculates a score for these novel objects to obtain a novelty score. This score reflects the novelty and uniqueness of the object.
[0057] The demand network identifies physiological and safety-related needs of visually impaired users in a semantic graph sequence through a demand attention mechanism, such as drinking water, toileting, tactile paving, and handrails, and calculates a demand score for each need. This score reflects the importance and urgency of the need to the user.
[0058] The environmental recommendation model integrates the scores from these four sub-networks to generate a comprehensive interest score, which is used to rank and recommend environmental objects. User terminal device 3 recommends environmental objects to the visually impaired user based on the recommendation list. The user can perceive environmental information through multimodal interaction, including audio guidance and haptic feedback. Haptic feedback is used to update the environmental recommendation model, enabling personalized recommendations.
[0059] As a preferred embodiment, the environment recommendation model also includes a user interaction adapter; the user interaction adapter is used to dynamically adjust the weights of background score, aesthetic score, novelty score and need score based on tactile feedback using a maximum likelihood estimation algorithm.
[0060] In this embodiment, the user interaction adapter is a key component of the environment recommendation model. It is responsible for learning user preferences through continuous interaction and dynamically adjusting the weights of environmental object ratings based on the user's tactile feedback. Specifically, the user interaction adapter uses the Maximum Likelihood Estimation (MLE) algorithm to dynamically adjust the weights of background rating, aesthetic rating, novelty rating, and need rating. The tactile feedback of visually impaired users is recorded and used as input to the MLE algorithm. For example, if a user touches an environmental object on the exploration interface for a longer period of time or touches it multiple times, it may indicate that the object has a high aesthetic or novelty level for the user.
[0061] Overall interest rating: , in, Weighting for aesthetic scores The weighting of the novelty rating, As the weight of the demand score, The weighting of the background score.
[0062] The user interaction adapter adjusts the weight of the corresponding ratings based on this feedback ("like / dislike" for each object). For example, if users show more interest in novelty objects, the user interaction adapter increases the weight of the novelty rating. This may also reduce the weight of the background score. This dynamic adjustment ensures that the recommendation list better reflects users' personal preferences and needs. Objects receiving the highest overall interest score (i.e., higher aesthetic score, novelty score, need score, and lower background score) are assigned a higher total score, thus gaining higher priority and ranking in the recommendation list. This MLE iteration allows the guided tour plan to be continuously updated during the testing phase, with the weight parameters reflecting users' personalized interests being optimized in real time. This process can be viewed as a crucial parameter in defining the priority of scene graph nodes through implicit user choices, while the environment recommendation model attempts to estimate the true value of these weights.
[0063] In this way, a personalized recommendation list can be generated for each user, where the order of environmental objects reflects the user's preferences for different environmental features. This not only improves user satisfaction but also enhances their interactive experience with the environment. The introduction of the user interaction adapter enables it to adaptively learn and predict user behavior patterns, thereby providing more accurate and personalized services. This dynamic adjustment mechanism allows the model to continuously optimize over time to adapt to users' changing needs and preferences.
[0064] The following describes the multimodal interactive sensing device for dynamic panoramic environment information provided by the present invention. The multimodal interactive sensing device for dynamic panoramic environment information described below and the multimodal interactive sensing method for dynamic panoramic environment information described above can be referred to in correspondence.
[0065] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a multimodal interactive sensing device for dynamic panoramic environmental information provided by the present invention.
[0066] This invention also provides a multimodal interactive sensing device for dynamic panoramic environment information, applied to a user terminal device 3. The device includes: a recommendation list acquisition module 201, used to acquire a recommendation list for environmental object perception; and an environmental object recommendation module 202, used to recommend environmental objects to visually impaired users through multimodal interaction based on the recommendation list, so as to assist visually impaired users in environmental perception from multiple dimensions. The multimodal interaction includes audio guidance and tactile feedback; the multiple dimensions include exploration, understanding, recall, and social interaction dimensions. The recommendation list is a list generated by the environment recommendation model based on the dynamic panoramic environment video, which sorts several environmental objects according to a comprehensive interest score. The comprehensive interest score is a score that comprehensively considers the aesthetic, novelty, and demand factors of the visually impaired user. The environment recommendation model is used to iteratively update the association weights of aesthetic, novelty, and demand factors based on tactile feedback, so as to iteratively update the comprehensive interest score and the recommendation list.
[0067] The following describes the multimodal interactive perception system for dynamic panoramic environment information provided by the present invention. The multimodal interactive perception system for dynamic panoramic environment information described below and the multimodal interactive perception method for dynamic panoramic environment information described above can be referred to in correspondence.
[0068] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of a multimodal interactive sensing system for dynamic panoramic environmental information provided by the present invention.
[0069] This invention also provides a multimodal interactive perception system for dynamic panoramic environmental information, comprising: a data acquisition device 1 for capturing dynamic panoramic environmental video and transmitting the dynamic panoramic environmental video to a server 2; a server 2 for running an environment recommendation model based on the dynamic panoramic environmental video to generate a recommendation list; the recommendation list is a list of several environmental objects sorted according to a comprehensive interest score; the comprehensive interest score is a score that integrates the aesthetic, novelty, and demand factors of visually impaired users; the environment recommendation model is used to iteratively update the association weights of aesthetic, novelty, and demand factors based on tactile feedback, so as to iteratively update the comprehensive interest score and the recommendation list; and a user terminal device 3 for executing the above-described multimodal interactive perception method for dynamic panoramic environmental information.
[0070] The system of this invention can simulate the perspective selection mechanism when a human accompanies a visually impaired person on a guided tour, and decomposes the environmental perception requirements into three core dimensions: Aesthetic value perception: Capture aesthetically significant elements in unfamiliar scenes to help users develop an understanding of the characteristics of the environment.
[0071] Dynamic Novelty Capture: Identify novel information in a scene that arises from changes in time and space (such as temporary performances, special installations, etc.) to enhance the user's perception of the dynamic environment.
[0072] Demand warning signals: Real-time detection of warning information related to basic needs (such as drinking water points, accessible facilities, obstacles, etc.) to ensure user safety.
[0073] This invention employs a self-supervised graph masking semantic graph training scheme. It utilizes only the inherent structural information in unlabeled data, employing a masking mechanism to allow the algorithm to automatically learn how to parse semantic relationships within a scene, thus eliminating reliance on manual annotation. This training method not only reduces data acquisition costs but also enhances the algorithm's environmental generalization ability, providing visually impaired users with a continuously optimized interactive experience.
[0074] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 401, a communications interface 402, a memory 403, and a communication bus 404, wherein the processor 401, the communications interface 402, and the memory 403 communicate with each other through the communication bus 404. The processor 401 can call logical instructions in the memory 403 to execute a multimodal interactive perception method for dynamic panoramic environment information, applied to the user terminal device 3. This method includes: obtaining a recommendation list of perceived environmental objects; recommending environmental objects to visually impaired users through multimodal interaction based on the recommendation list, thereby assisting visually impaired users in environmental perception from multiple dimensions; the multimodal interaction includes audio guidance and tactile feedback; the multiple dimensions include exploration, understanding, recall, and social interaction; wherein, the recommendation list is a list generated by the environment recommendation model based on the dynamic panoramic environment video, sorting several environmental objects according to a comprehensive interest score; the comprehensive interest score is a score that comprehensively considers the aesthetic, novelty, and needs factors of the visually impaired user; the environment recommendation model is used to iteratively update the association weights of aesthetic, novelty, and needs factors based on tactile feedback, thereby iteratively updating the comprehensive interest score and the recommendation list.
[0075] Furthermore, the logical instructions in the aforementioned memory 403 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server 2, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0076] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal interactive perception method for dynamic panoramic environment information provided by the above methods, applied to a user terminal device 3. The method includes: obtaining a recommendation list for environmental object perception; recommending environmental objects to visually impaired users through multimodal interaction based on the recommendation list, so as to assist visually impaired users in environmental perception from multiple dimensions; the multimodal interaction includes audio guidance and tactile hierarchical feedback; the multiple dimensions include exploration, understanding, recall, and social interaction dimensions; wherein, the recommendation list is a list generated by the environment recommendation model based on the dynamic panoramic environment video, which sorts several environmental objects according to a comprehensive interest score; the comprehensive interest score is a score that comprehensively considers the aesthetics, novelty, and needs of visually impaired users; the environment recommendation model is used to iteratively update the association weights of aesthetics, novelty, and needs based on tactile hierarchical feedback, so as to iteratively update the comprehensive interest score and the recommendation list.
[0077] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a multimodal interactive perception method for dynamic panoramic environmental information provided by the methods described above, applied to a user terminal device 3. The method includes: obtaining a recommendation list of environmental objects for perception; recommending environmental objects to a visually impaired user through multimodal interaction based on the recommendation list, so as to assist the visually impaired user in environmental perception from multiple dimensions; the multimodal interaction includes audio guidance and tactile feedback; the multiple dimensions include exploration, understanding, recall, and social interaction dimensions; wherein, the recommendation list is a list generated by an environmental recommendation model based on a dynamic panoramic environmental video, which sorts several environmental objects according to a comprehensive interest score; the comprehensive interest score is a score that comprehensively considers the aesthetic, novelty, and demand factors of the visually impaired user; the environmental recommendation model is used to iteratively update the association weights of aesthetic, novelty, and demand factors based on tactile feedback, so as to iteratively update the comprehensive interest score and the recommendation list.
[0078] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal interactive sensing method for dynamic panoramic environmental information, characterized in that, Applied to user terminal equipment, the method includes: Obtain an environmental object-aware recommendation list; Based on the recommendation list, environmental objects are recommended to visually impaired users through multimodal interaction to assist them in environmental perception from multiple dimensions; the multimodal interaction includes audio guidance and tactile feedback; the multiple dimensions include exploration, understanding, recall and social interaction. The recommended list is a list generated by the environmental recommendation model based on dynamic panoramic environmental videos, which sorts several environmental objects according to a comprehensive interest score. The comprehensive interest score is a score that combines the aesthetic, novelty, and demand factors of visually impaired users. The environmental recommendation model is used to iteratively update the association weights of aesthetic, novelty, and demand factors based on the tactile feedback, so as to iteratively update the comprehensive interest score and the recommended list. When the visually impaired user explores the environment, information about recommended environmental objects is described to the visually impaired user via audio based on the recommendation list; the information about the recommended environmental objects includes the name, location, and attributes of the recommended environmental objects. The exploratory actions of the visually impaired user are determined; the exploratory actions are used to characterize the visually impaired user's preference information for the currently recommended environmental objects; the preference information includes likes and dislikes. The recommendation list is dynamically adjusted based on the exploration actions. When the visually impaired user understands the environment, based on the recommendation list, a two-layer hierarchical scene interaction architecture is adopted to form several environmental objects in a main layer graph structure and a sub-layer graph structure; the nodes of the main layer graph structure are coarse-grained environmental objects, and the edges of the main layer graph structure are the relationships between the coarse-grained environmental objects; the nodes of the sub-layer graph structure are fine-grained environmental objects, and the edges of the sub-layer graph structure are the relationships between the fine-grained environmental objects. Based on the recommendation list, information about coarse-grained environmental objects in the main layer graph structure is described to the visually impaired user via audio. Determine the understanding actions of the visually impaired user; the understanding actions are used to characterize the visually impaired user's need to continue understanding information about finer-grained environmental objects under the currently recommended coarse-grained environmental objects; Based on the understanding action, the information of the fine-grained environment object in the sub-layer graph structure corresponding to the currently recommended coarse-grained environment object is described to the visually impaired user via audio.
2. The multimodal interactive perception method for dynamic panoramic environmental information according to claim 1, characterized in that, When the visually impaired user recalls their environment, the recall action of the visually impaired user is determined; the recall action is used to characterize the visually impaired user's need to recall historical environmental objects. Based on the recall action, historical multimodal interactive environmental perception information is retrieved, and information about the historical environmental objects is described to the visually impaired user via audio.
3. The multimodal interactive perception method for dynamic panoramic environmental information according to claim 1, characterized in that, When the visually impaired user engages in environmental social interaction, the social interaction action of the visually impaired user is determined; the social interaction action is used to characterize the current visually impaired user's need to share information about its historical environmental objects with another visually impaired user. Based on the social interaction actions, historical multimodal interactive environment perception information is retrieved and sent to the receiving device of the other visually impaired user, so that the current visually impaired user and the other visually impaired user can share the information of the historical environment objects.
4. The multimodal interactive sensing method for dynamic panoramic environment information according to any one of claims 1 to 3, characterized in that, Also includes: The dynamic panoramic environment video is converted into a semantic graph sequence using a scene graph generation algorithm, and the semantic graph sequence is then input into the environment recommendation model. The environment recommendation model is trained based on environment video training samples, using graph mask self-supervised learning and multimodal attention mechanism; The environmental recommendation model includes a background network, an aesthetic network, a novelty network, and a demand network. The background network is used to identify background objects in the semantic graph sequence through a background attention mechanism and to calculate the score of the background object to obtain a background score. The aesthetic network is used to identify objects of interest that continuously appear in the semantic graph sequence within a preset time period through an aesthetic attention mechanism, and to calculate the score of the objects of interest to obtain an aesthetic score. The novelty network is used to identify novelty objects that appear for the first time in the semantic graph sequence compared to the previous frame through a novelty attention mechanism, and to calculate the score of the novelty object to obtain a novelty score. The demand network is used to identify physiological and safety-related demand objects for the visually impaired user in the semantic graph sequence through a demand attention mechanism, and to calculate the score of the demand object to obtain a demand score. The comprehensive interest score is determined by combining the background score, the aesthetic score, the novelty score, and the need score.
5. The multimodal interactive perception method for dynamic panoramic environment information according to claim 4, characterized in that, The environment recommendation model also includes a user interaction adapter; The user interaction adapter is used to dynamically adjust the weights of the background score, the aesthetic score, the novelty score, and the need score based on the tactile feedback using a maximum likelihood estimation algorithm.
6. A multimodal interactive sensing device for dynamic panoramic environmental information, characterized in that, The device is applied to user terminal equipment and includes: The recommendation list retrieval module is used to obtain a recommendation list that is aware of environmental objects; An environmental object recommendation module is used to recommend environmental objects to visually impaired users through multimodal interaction based on the recommendation list, so as to assist the visually impaired users in environmental perception from multiple dimensions; the multimodal interaction includes audio guidance and tactile feedback; the multiple dimensions include exploration, understanding, recall and social interaction dimensions; The recommended list is a list generated by the environmental recommendation model based on dynamic panoramic environmental videos, which sorts several environmental objects according to a comprehensive interest score. The comprehensive interest score is a score that combines the aesthetic, novelty, and demand factors of visually impaired users. The environmental recommendation model is used to iteratively update the association weights of aesthetic, novelty, and demand factors based on the tactile feedback, so as to iteratively update the comprehensive interest score and the recommended list. When the visually impaired user explores the environment, information about recommended environmental objects is described to the visually impaired user via audio based on the recommendation list; the information about the recommended environmental objects includes the name, location, and attributes of the recommended environmental objects. The exploratory actions of the visually impaired user are determined; the exploratory actions are used to characterize the visually impaired user's preference information for the currently recommended environmental objects; the preference information includes likes and dislikes. The recommendation list is dynamically adjusted based on the exploration actions. When the visually impaired user understands the environment, based on the recommendation list, a two-layer hierarchical scene interaction architecture is adopted to form several environmental objects in a main layer graph structure and a sub-layer graph structure; the nodes of the main layer graph structure are coarse-grained environmental objects, and the edges of the main layer graph structure are the relationships between the coarse-grained environmental objects; the nodes of the sub-layer graph structure are fine-grained environmental objects, and the edges of the sub-layer graph structure are the relationships between the fine-grained environmental objects. Based on the recommendation list, information about coarse-grained environmental objects in the main layer graph structure is described to the visually impaired user via audio. Determine the understanding actions of the visually impaired user; the understanding actions are used to characterize the visually impaired user's need to continue understanding information about finer-grained environmental objects under the currently recommended coarse-grained environmental objects; Based on the understanding action, the information of the fine-grained environment object in the sub-layer graph structure corresponding to the currently recommended coarse-grained environment object is described to the visually impaired user via audio.
7. A multimodal interactive sensing system for dynamic panoramic environmental information, characterized in that, include: Acquisition equipment is used to capture dynamic panoramic environment video and transmit the dynamic panoramic environment video to a server; The server is used to generate a recommendation list based on a dynamic panoramic environment video running environment recommendation model; the recommendation list is a list of several environmental objects sorted according to a comprehensive interest score; the comprehensive interest score is a score that combines the aesthetics, novelty and needs of visually impaired users; the environment recommendation model is used to iteratively update the association weights of aesthetics, novelty and needs based on the tactile feedback, so as to iteratively update the comprehensive interest score and the recommendation list; The user terminal device is used to execute the multimodal interactive perception method for dynamic panoramic environment information as described in any one of claims 1 to 5.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal interactive perception method for dynamic panoramic environment information as described in any one of claims 1 to 5.