Indoor navigation method, equipment and program product

By using a multi-task learning system based on transformer architecture and virtual human interaction technology in indoor navigation, the problems of poor indoor navigation user experience and insufficient navigation accuracy in the existing technology are solved, and a more accurate and natural navigation experience is achieved.

CN120043534APending Publication Date: 2025-05-27SHENZHEN MEITUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332748.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has poor user experience in indoor navigation, can only provide static path guidance, and the navigation accuracy is insufficient.

Method used

A multi-task learning system based on transformer architecture analyzes the navigation requirements of users' voice descriptions, combines the building's topological structure library to generate navigation routes, and updates locations in real time through multiple rounds of dialogue between virtual people and users, and issues navigation voice commands.

Benefits of technology

Providing accurate indoor navigation without reliable hardware positioning support improves navigation accuracy and user experience, making the navigation process more natural and intuitive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120043534A_ABST
    Figure CN120043534A_ABST
Patent Text Reader

Abstract

The invention provides an indoor navigation method, equipment and a program product, and relates to the technical field of navigation. The method comprises the following steps: acquiring a navigation demand of a user, analyzing the navigation demand by utilizing a transformer architecture-based multi-task learning system, and determining a current position and a destination position which need to be navigated; generating a navigation route based on the current location, the destination location and the topological structure library of the building, the navigation route including a plurality of points of interest between the current location and the destination location; the current position of the user is updated in real time through multiple rounds of conversations between the virtual human and the user; under the condition that the current position does not deviate from the navigation route, in response to updating of the current position, the virtual human is controlled to send a navigation voice instruction, and the navigation voice instruction is used for guiding the user to go to the next interest point of the navigation route from the current position. According to the embodiment of the invention, the experience of obtaining navigation help by a video call with a real shopping guide can be simulated, and more natural and visual navigation experience is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of indoor navigation technology, and in particular, to an indoor navigation method, device, and program product. Background Art

[0002] With the rapid development of mobile devices and positioning technologies, navigation systems have become an indispensable part of people's daily lives. Currently, most navigation applications and services focus on route planning and navigation in outdoor environments, such as road navigation services based on the Global Positioning System (GPS), which can provide users with accurate route guidance from a starting point to a destination and offer real-time updates when traffic conditions change.

[0003] However, when it comes to indoor environments, existing navigation solutions are relatively insufficient. Different from outdoor spaces, indoor structures are more complex and variable, including but not limited to the interiors of large buildings such as shopping malls, airports, hospitals, and museums. These places usually contain multiple floors, narrow passages, and complex layouts, making it impossible for traditional GPS signals to effectively penetrate buildings, resulting in a significant decrease or even complete failure of positioning accuracy. Therefore, there are few navigation solutions for indoor scenarios, and the existing solutions have a poor user experience, can only provide static route guidance, and have insufficient navigation accuracy.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The present disclosure provides an indoor navigation method, device, and program product, which at least to a certain extent overcome the problems in the related art that the indoor navigation solution has a poor user experience, can only provide static route guidance, and has insufficient navigation accuracy.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.

[0007] According to one aspect of the present disclosure, an indoor navigation method is provided, including: obtaining a user's navigation requirement, analyzing the navigation requirement by using a multi-task learning system based on the transformer architecture to determine a current location and a destination location that need navigation, the current location and the destination location being inside the same building, and the navigation requirement including one or more segments of voice data; generating a navigation route based on the current location, the destination location, and a topological structure library of the building, the navigation route including a plurality of points of interest between the current location and the destination location, and the topological structure library being a database storing information on the internal space layout of the building; updating the user's current location in real time through multiple rounds of conversations between a virtual human and the user; and when the current location does not deviate from the navigation route, in response to the update of the current location, controlling the virtual human to issue a navigation voice instruction for guiding the user to go from the current location to the next point of interest on the navigation route.

[0008] According to another aspect of the present disclosure, an indoor navigation device is provided, including a requirement acquisition module, a route generation module, a location update module, and a navigation interaction module.

[0009] The requirement acquisition module is configured to obtain a user's navigation requirement, analyze the navigation requirement by using a multi-task learning system based on the transformer architecture to determine a current location and a destination location that need navigation, the current location and the destination location being inside the same building, and the navigation requirement including one or more segments of voice data;

[0010] The route generation module is configured to generate a navigation route based on the current location, the destination location, and a topological structure library of the building, the navigation route including a plurality of points of interest between the current location and the destination location, and the topological structure library being a database storing information on the internal space layout of the building;

[0011] The location update module is configured to update the user's current location in real time through multiple rounds of conversations between a virtual human and the user;

[0012] The navigation interaction module is configured to, when the current location does not deviate from the navigation route, in response to the update of the current location, control the virtual human to issue a navigation voice instruction for guiding the user to go from the current location to the next point of interest on the navigation route.

[0013] According to still another aspect of the present disclosure, an electronic device is provided, including: a memory for storing instructions; and a processor for calling the instructions stored in the memory to implement the above indoor navigation method.

[0014] According to still another aspect of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above indoor navigation method is implemented.

[0015] According to another aspect of the present disclosure, there is provided a computer program product. The computer program product stores instructions which, when executed by a computer, cause the computer to implement the indoor navigation method described above.

[0016] According to another aspect of the present disclosure, there is provided a chip, including at least one processor and an interface;

[0017] The interface is configured to provide program instructions or data for at least one processor;

[0018] At least one processor is configured to execute program instructions to implement the indoor navigation method described above.

[0019] The indoor navigation method, device and program product provided by the embodiments of the present disclosure are different from the traditional ways of obtaining the user's location by relying on hardware devices such as GPS, Wi-Fi, or Bluetooth beacons. The present disclosure uses a multi-task learning system based on the transformer architecture to parse the navigation requirements described in the user's speech to determine the location. This method can work without reliable hardware positioning support, especially suitable for scenarios with poor signals or where positioning services are not allowed; the pre-set topological structure library contains the spatial layout information inside the building. Since it is based on the known building structure rather than dynamic positioning data, theoretically, some problems caused by positioning errors can be avoided, such as misjudgment between multiple floors or inaccurate positioning inside walls, and thus the best path from a current location to a destination location can be accurately calculated to obtain a more suitable navigation route for the user; in addition, during the navigation process, the virtual human continuously communicates with the user, simulating the experience of obtaining navigation assistance through a video call with a real salesperson, providing a more natural and intuitive navigation experience.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0022] Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 Show a flowchart of an indoor navigation method in an embodiment of the present disclosure;

[0024] Figure 2 Show a schematic diagram of a virtual human interaction page in an embodiment of the present disclosure;

[0025] Figure 3 Flow chart showing the determination of the navigation position in an embodiment of the present disclosure

[0026] Figure 4 Schematic diagram of another virtual human interaction page in an embodiment of the present disclosure;

[0027] Figure 5 Schematic diagram of yet another virtual human interaction page in an embodiment of the present disclosure;

[0028] Figure 6 Flow chart showing the generation of the navigation route in an embodiment of the present disclosure

[0029] Figure 7 Flow chart of yet another indoor navigation method in an embodiment of the present disclosure;

[0030] Figure 8 Flow chart of still another indoor navigation method in an embodiment of the present disclosure;

[0031] Figure 9 Schematic diagram of an indoor navigation device in an embodiment of the present disclosure;

[0032] Figure 10 Block diagram of the structure of an electronic device in an embodiment of the present disclosure. Detailed implementation manners

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Usually, the components of the embodiments of the present disclosure described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the claimed present disclosure, but merely represents the selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0034] Traditional indoor positioning solutions, such as methods based on technologies such as Wi-Fi, Bluetooth, geomagnetism, and UWB, generally have problems of insufficient generality, easy decline in accuracy affected by the environment, and high maintenance costs, which limit their wide adoption in practical applications.

[0035] Bluetooth technology requires the deployment of numerous beacon devices in indoor environments, which not only increases the hardware and maintenance costs, but may also encounter signal instability in complex environments, thus affecting the accuracy of navigation. Augmented Reality (AR) technology can provide intuitive visual navigation guidance, but its implementation relies on complex hardware such as high-performance cameras and processors, and requires users to continuously hold the mobile device and align it with the surrounding environment, which undoubtedly reduces the quality of the user experience.

[0036] Traditional indoor navigation systems lack intelligent interaction functions and are limited to providing fixed path guidance, unable to make flexible adjustments based on the user's real-time location and personalized needs.

[0037] In the embodiments of the present disclosure, through the form of dialogue with a virtual human, the multi-task learning system is combined to clarify the user's needs in real time, and real-time communication is carried out during the user's movement to provide dynamic navigation suggestions, simulating the experience of obtaining navigation assistance through a video call with a real salesperson, and providing a more natural and intuitive navigation experience.

[0038] It should be noted that the data involved in the embodiments of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0039] The indoor navigation method of the embodiments of the present disclosure can be applied to an electronic device. The execution subject of the indoor navigation method can be at least one of user terminals such as mobile phones, tablets, wearable devices, etc. that can be configured to execute the indoor navigation method provided by the embodiments of the present disclosure, or the execution subject of the method can also be the client itself that can execute the method.

[0040] The following will describe this exemplary embodiment in detail with reference to the accompanying drawings and embodiments.

[0041] Figure 1 The flowchart of an indoor navigation method in the embodiments of the present disclosure is shown. As Figure 1 shown, the indoor navigation method provided in the embodiments of the present disclosure includes S101 - S104.

[0042] In S101, the navigation requirements of the user are obtained, and the multi-task learning system based on the transformer architecture is used to analyze the navigation requirements to determine the current location and the destination location that need to be navigated. The current location and the destination location are inside the same building, and the navigation requirements include one or more segments of voice data.

[0043] The execution subject of the indoor navigation method can be an electronic device or a client in the electronic device. The client can include a virtual human interaction page. The virtual human interaction page can be as Figure 2As shown, the page may include a voice input control 201. In one example, the execution subject of the indoor navigation method may be a terminal. The terminal may collect sound signals (including the user's voice) when the user holds down the voice input control 201 and stop collecting when the user releases the voice input control 201, thereby obtaining a segment of voice data. The user may input multiple times, and thus the terminal may obtain multiple segments of voice data.

[0044] In some embodiments, the above-mentioned one or more segments of voice data may be converted into text data by using voice recognition technology, and then a multi-task learning system based on the Transformer architecture may be adopted to parse the user's intention based on the text data and determine the current position and destination position that need to be navigated.

[0045] The application scenario of the above navigation method is an indoor navigation scenario, that is, the current position and the destination position are located inside the same building. The building may be a large building such as a shopping mall, an airport, a hospital, a museum, etc. The above building may include multiple floors, and the current position and the destination position may be located on different floors.

[0046] In S102, a navigation route is generated based on the current position, the destination position, and the topological structure library of the building. The navigation route includes multiple points of interest between the current position and the destination position. The topological structure library is a database storing the internal space layout information of the building.

[0047] A point of interest (POI) is a location that a user may be interested in or need to visit. A point of interest may also be referred to as a key node of the navigation route. In some embodiments, a point of interest may be a specific location within a building, which may be a functional area, a facility, or any place meaningful to the user. For example, in a shopping mall, the points of interest may be stores, restrooms, elevators, stairs, service desks, etc.; in a hospital, they may be consulting rooms, pharmacies, waiting areas, examination rooms, etc. When generating a navigation route, points of interest are taken into consideration as important reference points. They are not only intermediate stops for the user from the current position to the destination but also help the user better understand and remember the path.

[0048] In some embodiments, the information of points of interest may be dynamically managed and updated. For example, when a certain store closes or changes its business content, the system should promptly reflect this change. Furthermore, when the internal layout of the building changes, the points of interest can be updated in a timely manner. In addition, users may be encouraged to feedback new points of interest or report status changes of existing points of interest to further improve the accuracy and practicality of the navigation route.

[0049] The topological structure library is a database that stores information about the internal spatial layout of a building, including data on key locations such as floor plans, rooms, corridors, stairs, elevators, etc., as well as the connection relationships between them. Each point of interest (such as a coffee shop, restroom, exit, etc.) should have a clear coordinate identification.

[0050] Based on the current location, destination location, and the topological structure library of the building, a path planning algorithm is called to find the shortest or optimal path, obtaining a navigation route. Among them, the path planning algorithm can be the Dijkstra algorithm, A (A-Star) search algorithm, D Lite algorithm, Floyd-Warshall algorithm, graph traversal algorithm, etc.

[0051] As an example, the above S102 can be to retrieve the shortest path in the topological structure library and display the walking navigation route on the user device, including all the points of interest (POIs) passed through.

[0052] In some embodiments, when generating the navigation route, not only the distance factor is considered, but also factors such as time cost and crowd density can be comprehensively evaluated to provide more optimized route suggestions.

[0053] The embodiments of the present disclosure are based on a known topological structure library rather than dynamic positioning data. Theoretically, some problems caused by positioning errors can be avoided, such as misjudgment between multiple floors or inaccurate positioning within walls. Furthermore, the best path from a current location to a destination location can be accurately calculated, obtaining a navigation route more suitable for users.

[0054] In S103, the current location of the user is updated in real time through multiple rounds of conversations between the virtual human and the user.

[0055] The embodiments of the present disclosure adopt a lightweight 3D rendering engine based on WebGL to implement 3D modeling of virtual humans, bypassing the complex processes of traditional 3D modeling, and being able to directly generate the virtual human image through text description.

[0056] In some embodiments, a component library can also be preset, covering different facial features, hairstyles, clothing styles, limb movements, etc. Furthermore, users can quickly create unique virtual images by selecting or mixing these preset components, freely customizing the appearance, clothing, limb movements, etc. of the virtual human.

[0057] In some embodiments, the present disclosure can also accurately synchronize lip movements, limb movements, and facial expressions, achieving a real-time rendering frame rate ≥ 30fps on the mobile side and a memory occupancy rate < 15% of the available memory of the terminal.

[0058] Throughout the navigation process, the virtual human maintains contact with the user through voice interaction (multi-round dialogue) to confirm the user's location and status. When the user reports that they have reached a specific point of interest, the system can immediately update its current location information. For example, if the user says "I'm at the elevator", the system can set the elevator as the user's current location.

[0059] Different from the traditional method of relying on hardware devices such as GPS, Wi-Fi, or Bluetooth beacons to obtain the user's location, the indoor navigation method of the embodiments of the present disclosure uses a multi-task learning system to analyze the navigation requirements described by the user's voice to determine the location. This method can work without reliable hardware positioning support, especially suitable for scenarios with poor signal or where positioning services are not allowed.

[0060] In S104, when the current location does not deviate from the navigation route, in response to the update of the current location, control the virtual human to issue a navigation voice instruction, and the navigation voice instruction is used to guide the user from the current location to the next point of interest on the navigation route.

[0061] When the current location does not deviate from the navigation route, the embodiments of the present disclosure can intelligently respond to the update of the user's current location and timely issue clear and accurate navigation voice instructions through the virtual human to guide the user to the next point of interest on the navigation route.

[0062] Whenever new location data is received, the system compares the user's current location with the preset navigation route to confirm whether the user is still on the correct direction of travel. In some embodiments, when the user's current location is within a predetermined distance (e.g., 5 meters or less) of the next point of interest, it is considered that the user is about to reach that point of interest and is ready to issue the next navigation voice instruction. In some embodiments, the predetermined distance for triggering the next navigation voice instruction can be determined according to the user's walking speed. The embodiments of the present disclosure can determine the user's location through the voice interaction between the virtual human and the user, and then calculate the user's walking speed.

[0063] During the navigation process, the virtual human of the embodiments of the present disclosure continuously communicates with the user, simulating the experience of obtaining navigation assistance through a video call with a real shopping guide, providing a more natural and intuitive navigation experience.

[0064] In some embodiments, as Figure 3 shown, the above S101 may include steps S301 - S304.

[0065] S301, obtain the user's navigation requirements;

[0066] S302, use a multi-task learning system based on the transformer architecture to analyze the navigation requirements to obtain an analysis result;

[0067] S303. Confirm whether the current location and the destination location in the analysis result are accurate through multiple rounds of conversations between the virtual human and the user.

[0068] In case of inaccuracy, transfer to S301 and repeat the above steps until the user confirms that the current location and the destination location in the analysis result are accurate, and then transfer to S304 to obtain the current location and the destination location that the user needs to navigate.

[0069] As Figure 4 shown, each segment of voice data input by the user can be converted into text data and displayed on the virtual human interaction page.

[0070] In some embodiments, to obtain the navigation requirements of the user, a multi-task learning system based on the transformer architecture is used to analyze the navigation requirements, and the current location and the destination location that need to be navigated are determined. It can be through the conversation between the virtual human and the user to obtain the navigation requirements of the user, and use the multi-task learning system based on the transformer architecture to convert the location described in the natural language of the user in the navigation requirements into a spatial coordinate system to obtain the analysis result; through multiple rounds of conversations between the virtual human and the user, confirm whether the current location and the destination location in the analysis result are accurate. In case of inaccuracy, have a conversation between the virtual human and the user again, and use the multi-task learning system to analyze until the user confirms that the current location and the destination location in the analysis result are accurate.

[0071] In the above solution, both the current location and the destination location are obtained through the analysis of the voice interaction content between the virtual human and the user and are recognized by the user. Compared with the method of obtaining positioning through hardware devices such as GPS, Wi-Fi, and Bluetooth beacons, it is more similar to the way of a real shopping guide pointing the way, providing a better user experience and being more suitable for indoor scenarios with poor positioning signals.

[0072] In some embodiments, through multiple rounds of conversations between the virtual human and the user, the current location of the user can be updated in real time, which may include combining the walking speed of the user to obtain the current estimated location of the user, triggering a conversation between the virtual human and the user before the user reaches the next point of interest to determine whether the estimated location is accurate, and using the multi-task learning system to analyze the conversation to update the current location of the user.

[0073] The virtual human can communicate with the user by voice, ask about the user's location, and can also calculate the speed at which the user is moving based on the user's answer, and thus can also predict the user's location.

[0074] In the embodiments of the present disclosure, through the semantic coordinate mapping mechanism in multiple rounds of conversations, the user's location description (such as "in front of HeyTea") is converted into a spatial coordinate system. During the conversation, the active inquiry is triggered by estimating the user's walking speed, so that the navigation service is transformed from one-way (the user puts forward the navigation requirement and the system gives the navigation plan) to two-way service (the user puts forward the navigation requirement, the system gives the navigation plan, the system follows up the user's navigation progress and adjusts the navigation plan in real time).

[0075] In some embodiments, after generating the navigation route based on the current location, the destination location and the topological structure library of the building in S102 above, the navigation route can also be displayed on the virtual human interaction page. Among them, the navigation route can be a navigation route in text form, and / or a navigation route in map form. As an example, the display page for displaying the navigation route can be as Figure 5 shown, which may include a navigation route 501 in text form and a navigation route 502 in map form.

[0076] In some embodiments, as Figure 6 shown, generating the navigation route based on the current location, the destination location and the topological structure library of the building in S102 above may include S601-S602.

[0077] In S601, on the virtual human interaction page, multiple routes to be selected are displayed, and the multiple routes to be selected are generated based on the current location, the destination location and the topological structure library of the building;

[0078] In S602, in response to the selection operation of the navigation route among the multiple routes to be selected, the navigation route is displayed.

[0079] In some embodiments, when the user is not satisfied with the generated navigation route, the user can continue to interact with the virtual human and put forward new requirement information. The embodiments of the present disclosure can determine the route suggestion of the user for the navigation route through the voice interaction between the virtual human and the user; and generate a new navigation route based on the current location, the destination location, the topological structure library of the building and the route suggestion.

[0080] In some embodiments, the above route suggestion can be the special requirement of the user, such as choosing a path with escalators instead of stairs, or providing a barrier-free passage for users carrying large luggage, etc.

[0081] In some embodiments, the user can also propose the above route suggestions in step S101. Furthermore, S101 can be to obtain the navigation requirements input by the user on the virtual human interaction page, analyze the navigation requirements using a multi-task learning system, and determine the current location, destination location, and route suggestions that need navigation. The above S102 can be to generate a navigation route based on the current location, destination location, topological structure library of the building, and route suggestions.

[0082] Figure 7 The flowchart of an indoor navigation method in an embodiment of the present disclosure is shown, as Figure 7 shown, the indoor navigation method provided in the embodiment of the present disclosure includes S701 - S706. Among them, S701 - S704 are the same as S101 - S104 in the above embodiment and will not be elaborated here.

[0083] In S705, when the current location deviates from the navigation route, a new navigation route is generated based on the updated current location, destination location, and topological structure library of the building;

[0084] In S706, control the virtual human to issue a navigation voice instruction, and the navigation voice instruction is used to guide the user to go from the updated current location to the next point of interest on the new navigation route.

[0085] The embodiment of the present disclosure introduces a virtual human as an interaction interface and communicates through voice data, making the navigation process more natural and smooth. The virtual human can update the position in real time according to the actual movement of the user. When the current location does not deviate from the navigation route, corresponding voice instructions are issued in response to the position update to ensure that the user is always in the correct traveling direction. If a deviation is detected, the route can be adjusted immediately or new instructions can be provided to maintain the effectiveness and accuracy of the navigation.

[0086] It should be noted that the multi-task learning system based on the transformer architecture can be deployed on the electronic device itself, that is, deployed in the execution entity of the present disclosure solution; the multi-task learning system based on the transformer architecture can also be deployed on the server side, and the execution entity of the present disclosure solution can be the client. Taking the example of the multi-task learning system based on the transformer architecture being deployed on the server side, the conversation between the virtual human and the user will be described below.

[0087] As Figure 8 shown, the conversation between the virtual human and the user includes S801 - S807.

[0088] In S801, the client obtains the voice data input by the user and the facial image of the user when the voice input is made.

[0089] In S802, the client sends voice data and facial images to the server.

[0090] In S803, the server analyzes the text semantics based on the voice data and recognizes the user's voice emotion based on the intonation of the voice data.

[0091] In S804, the server recognizes the user's visual emotion based on the facial image when the user makes a voice input.

[0092] In S805, based on the text semantics, the voice emotion, and the visual emotion, a reply voice for the virtual human to make a voice reply is generated, and the intonation and facial expression for the virtual human to make a voice reply are determined, and a first instruction is generated.

[0093] Emotion recognition based on voice and emotion detection based on facial images can capture the subtle emotional changes of users. For example, frowning indicates confusion or dissatisfaction. The virtual human can dynamically adjust its response tone and expression according to these clues, making the interaction process more natural and smooth, as if having a conversation with a real person.

[0094] In S806, the server sends the first instruction to the client.

[0095] In S807, in response to the first instruction from the server, the virtual human makes a voice reply.

[0096] The multi-modal fusion processor of the embodiments of the present disclosure integrates voice emotion recognition, visual emotion detection, and text semantic analysis, and dynamically adjusts the virtual human interaction strategy. For example, when detecting a frowning expression of the user, it adjusts the response tone or adjusts the navigation plan. The virtual human can more accurately understand the user's needs and emotional state, so as to provide a response and service that better fits the user's current situation, realizing a high degree of interaction and personalized communication between the virtual human and the user, which greatly enhances the user's interaction experience, making the reply of the virtual human not only limited to the information level, but also including the understanding and support at the emotional level.

[0097] In some embodiments, the above S802 sending voice data to the server may include: instantaneously encoding and fragmenting the voice data by using ASR (Automatic Speech Recognition) streaming processing technology; encapsulating the instantaneously encoded and fragmented data through the RTC (Real-Time Communication) protocol and sending it to the server.

[0098] ASR streaming processing allows for the encoding and processing of voice data while the user is speaking, rather than waiting until the entire voice input is complete before starting processing, which can significantly reduce the processing time and improve the response speed. Combining with the RTC protocol for data transmission ensures that voice data can be sent from the client to the server with the lowest possible latency. The embodiments of the present disclosure adopt a combination of ASR streaming processing technology and RTC protocol, which not only improves the speed and efficiency of voice data processing, but also greatly enhances the user experience, making real-time voice interaction more fluent, stable and reliable.

[0099] In some embodiments, during the navigation process, the user can interrupt and interject in real time through voice, text, etc. That is to say, when the virtual human makes a voice reply, the user can interrupt the virtual human's reply. Correspondingly, when the virtual human makes a voice reply, the present disclosure uses VAD technology to detect whether there is new voice activity or text input; if new voice activity or text input is detected, semantic integrity evaluation is started to obtain an evaluation result; when the interruption response window time is reached, the user's voice activity and / or text input within the interruption response window time are used to determine the user's intention, and it is determined whether interruption is needed based on the user's intention; when interruption is needed, the interruption point in the current context is determined according to the evaluation result; the voice reply being made by the virtual human is paused at the interruption point, and new voice activity or text input from the user is received.

[0100] The present disclosure first uses voice activity detection (VAD) technology to monitor whether there is new voice input. Once new voice activity is detected, semantic integrity evaluation of the current utterance will be immediately started, aiming to determine whether the current speaker's utterance has completed a logical unit or is in a state where it can be safely interrupted.

[0101] At the same time, the present disclosure monitors the situation within a period of time starting from the detection of a possible interruption intention according to a preset adjustable interruption response window time of 0.1 to 1.5 seconds. During this period, not only whether there is new voice input is concerned, but also the result of semantic integrity evaluation is combined to determine the best interruption timing.

[0102] Therefore, the evaluation of the interruption point actually includes two aspects of considerations:

[0103] Immediate semantic integrity evaluation: It is executed immediately when new voice activity is detected, and is used to determine whether the current utterance has reached a suitable pause point or end point.

[0104] Dynamic evaluation within the interruption response window: During this window time, continuous monitoring and comprehensive consideration of the changes in semantic integrity and user intention are carried out to determine the most suitable interruption timing.

[0105] The evaluation of breakpoints in the embodiments of the present disclosure is a dynamic process that changes with the update of information within the window period, capable of more accurately capturing the user's interruption intention while maintaining the fluency and naturalness of the conversation.

[0106] In some embodiments, the user can also dynamically adjust parameters such as interruption sensitivity, conversation rhythm, and tone settings according to specific needs to better adapt to different interaction scenarios. The time length of the interruption response window can be flexibly adjusted according to actual needs, usually set within the range of 0.1 to 1.5 seconds, during which interruption requests will be prepared for processing. For example, if set to 0.1 seconds, it is very sensitive to interruptions and can almost immediately respond to the user's interruption; while if set to 1.5 seconds, it will wait for a longer time to confirm whether there is indeed an interruption intention to reduce the possibility of misjudgment.

[0107] In some embodiments, before generating a navigation route based on the above-mentioned current location, destination location, and building topology library, when the dynamic attributes associated with the destination location meet the preset conditions, a virtual human can also communicate with the user to recommend a new destination location; wherein, the new destination location is determined according to the user's historical behavior; the dynamic attributes include business hours and / or foot traffic.

[0108] The embodiments of the present disclosure can also set up a semantic topology library to support the dynamic annotation and relationship reasoning of the dynamic attributes of indoor POIs (such as shops, elevators). For example, it can recommend a path combination of "coffee shop → fast food restaurant" based on the user's historical behavior and associate POI attributes (such as business hours, foot traffic), so as to give correct navigation suggestions (such as: The Starbucks you usually go to is temporarily closed today. There is also a Luckin Coffee on the second floor, which is also good. Do you consider changing to this store?).

[0109] In some embodiments, the above-mentioned indoor navigation method also determines that the user has reached the destination location through voice interaction between the virtual human and the user, and displays a service recommendation page. After the user arrives at the destination, the virtual human ends the navigation and can provide additional services or suggestions, such as "preferential recommendations near the destination" and "destination merchant activity notifications".

[0110] In the embodiments of the present disclosure, the virtual human can communicate with the user via voice, ask about the user's location, and calculate the user's traveling speed based on the user's answer, and thus can also predict the user's location. In some embodiments, the virtual human can ask the user via voice whether they have reached the destination. For example, if the destination is a cinema, the virtual human can ask, "Have you arrived at the cinema yet?" This can ensure the accuracy of the system's judgment. If it is confirmed that the user has reached the destination, the virtual human can display a personalized service recommendation page based on the user's historical behavior, preferences, and current environmental factors. This may include nearby restaurants, store promotions, entertainment facilities, etc. The content of the service recommendation page can also be dynamically adjusted according to real-time situations. For example, in a shopping mall, ongoing discount activities can be displayed; at an airport, boarding gate information and duty-free shop offers can be provided. In some embodiments, the above indoor navigation method may further include determining, through voice interaction between the virtual human and the user, that the user needs to go to a new destination location inside the building; and generating a new navigation route based on the user's current location, the new destination location, and the topological structure library of the building.

[0111] The solution of determining a new destination location through voice interaction between the virtual human and the user and generating a new navigation route based on the current location, the new destination location, and the topological structure library of the building greatly enhances the interactivity and intelligence level of the indoor navigation system, providing a more convenient and personalized navigation experience for users. Users can express new navigation needs at any time, and the embodiments of the present disclosure can respond immediately and re-plan the route, improving flexibility and practicality.

[0112] In some embodiments, before the above S101 of obtaining the navigation requirements input by the user on the virtual human interaction page and analyzing the navigation requirements using a multi-task learning system to determine the current location and destination location to be navigated, the above indoor navigation method may further include: when the navigation trigger condition is met, starting the virtual human interaction page and prompting the user to describe the navigation requirements via voice; where the navigation trigger condition includes detecting that the voice input by the user contains a preset keyword, or identifying that the user needs navigation using a multi-task learning system.

[0113] It should be noted that in the voice interaction process between the virtual human and the user in any of the above embodiments, the multi-task learning system can be called to understand the user's natural language and complete the interaction. In some embodiments, the above multi-task learning system based on the transformer architecture can be a large language model.

[0114] In the embodiments of the present disclosure, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0115] The term "and / or" in this disclosure is merely a correlative relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship.

[0116] Furthermore, although the steps of the methods in this disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result.

[0117] In some embodiments, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.

[0118] Based on the same inventive concept, an indoor navigation device is also provided in the embodiments of this disclosure, as described in the following embodiments. Since the principle of solving problems in the device embodiments is similar to that of the above method embodiments, the implementation of the device embodiments can refer to the implementation of the above method embodiments, and the repeated parts will not be elaborated.

[0119] Figure 9 The schematic diagram of an indoor navigation device in the embodiments of this disclosure is shown, as Figure 9 shown, the indoor navigation device includes a demand acquisition module 901, a route generation module 902, a position update module 903, and a navigation interaction module 904.

[0120] The demand acquisition module 901 is used to acquire the navigation demand of the user, analyze the navigation demand using a multi-task learning system based on the transformer architecture, determine the current position and the destination position that need to be navigated, the current position and the destination position are inside the same building, and the navigation demand includes one or more segments of voice data;

[0121] The route generation module 902 is used to generate a navigation route based on the current position, the destination position, and the topological structure library of the building. The navigation route includes multiple points of interest between the current position and the destination position, and the topological structure library is a database storing the internal space layout information of the building;

[0122] The position update module 903 is used to update the current position of the user in real time through multiple rounds of conversations between the virtual human and the user;

[0123] The navigation interaction module 904 is configured to, when the current location does not deviate from the navigation route, in response to the update of the current location, control the virtual human to issue a navigation voice instruction, and the navigation voice instruction is used to guide the user from the current location to the next point of interest on the navigation route.

[0124] In some embodiments, the requirement acquisition module 901 is configured to obtain the user's navigation requirement through the conversation between the virtual human and the user, convert the location described in the user's natural language in the navigation requirement into a spatial coordinate system by using a multi-task learning system based on the transformer architecture to obtain an analysis result; confirm whether the current location and the destination location in the analysis result are accurate through multiple rounds of conversations between the virtual human and the user, and in case of inaccuracy, conduct conversations between the virtual human and the user again and perform analysis by using the multi-task learning system until the user confirms that the current location and the destination location in the analysis result are accurate.

[0125] In some embodiments, the location update module 903 is configured to combine the user's walking speed to obtain the user's current estimated location, trigger a conversation between the virtual human and the user before the user reaches the next point of interest, determine whether the estimated location is accurate, and analyze the conversation by using the multi-task learning system to update the user's current location.

[0126] In some embodiments, the navigation interaction module 904 is further configured to, when the current location deviates from the navigation route, generate a new navigation route based on the updated current location, destination location, and the topological structure library of the building; control the virtual human to issue a navigation voice instruction, and the navigation voice instruction is used to guide the user from the updated current location to the next point of interest on the new navigation route.

[0127] In some embodiments, the indoor navigation device further includes a service recommendation module.

[0128] The service recommendation module is configured to determine that the user reaches the destination location through voice interaction between the virtual human and the user, and display a service recommendation page.

[0129] The concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependent relationships.

[0130] Regarding the indoor navigation device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the indoor navigation method, and will not be elaborated herein.

[0131] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory.

[0132] In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.

[0133] Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0134] The following refers to Figure 10 to describe the electronic device provided by the embodiments of the present disclosure. Figure 10 The displayed electronic device 1000 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0135] Figure 10 The schematic diagram of the architecture of an electronic device 1000 provided by the embodiments of the present disclosure is shown. As Figure 10 shown, the electronic device 1000 includes but is not limited to: at least one processor 1010 and at least one memory 1020.

[0136] The memory 1020 is used to store instructions.

[0137] In some embodiments, the memory 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202, and may further include a read-only storage unit (ROM) 10203.

[0138] In some embodiments, the memory 1020 may further include a program / utility 10204 having a set (at least one) of program modules 10205. Such program modules 10205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0139] In some embodiments, the memory 1020 may store an operating system. The operating system may be an operating system such as a Real Time eXecutive (RTX), LINUX, UNIX, WINDOWS, or OS X.

[0140] In some embodiments, data may also be stored in the memory 1020.

[0141] As an example, the processor 1010 may read data stored in the memory 1020. The data may be stored at the same storage address as the instruction, or the data may be stored at a different storage address from the instruction.

[0142] The processor 1010 is configured to call instructions stored in the memory 1020 to implement the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification. For example, the processor 1010 may execute the steps of the indoor navigation method embodiment described above.

[0143] It should be noted that the above-mentioned processor 1010 may be a general-purpose processor or a dedicated processor. The processor 1010 may include one or more processing cores, and the processor 1010 executes various functional applications and data processing by running instructions.

[0144] In some embodiments, the processor 1010 may include a central processing unit (CPU) and / or a baseband processor.

[0145] In some embodiments, the processor 1010 may determine an instruction according to the priority identifier and / or function category information carried in each control instruction.

[0146] In the present disclosure, the processor 1010 and the memory 1020 may be provided separately or integrated together.

[0147] As an example, the processor 1010 and the memory 1020 may be integrated on a single board or a system on chip (SOC).

[0148] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The electronic device 1000 may further include a bus 1030.

[0149] The bus 1030 may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures.

[0150] The electronic device 1000 can also communicate with one or more external devices 1040 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1050.

[0151] Moreover, the electronic device 1000 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1060.

[0152] As Figure 10 shown, the network adapter 1060 communicates with other modules of the electronic device 1000 through the bus 1030.

[0153] It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0154] It can be understood that the structure illustrated in the embodiments of the present disclosure does not constitute a specific limitation on the electronic device 1000. In other embodiments of the present disclosure, the electronic device 1000 may include more or fewer components than Figure 10 shown, or combine certain components, or split certain components, or have different component arrangements. Figure 10 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0155] The present disclosure also provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the indoor navigation method described in the above method embodiments is implemented.

[0156] In the embodiments of the present disclosure, the computer-readable storage medium is a medium that can send, propagate, or transmit computer instructions for use by or in connection with an instruction execution system, apparatus, or device.

[0157] As an example, the computer-readable storage medium is a non-volatile storage medium.

[0158] In some embodiments, more specific examples of the computer-readable storage medium in the present disclosure may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, USB flash drives, portable hard disks, or any suitable combination of the above.

[0159] In an embodiment of the present disclosure, the computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer instructions (readable program code).

[0160] Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.

[0161] In some examples, the computing instructions included on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0162] The embodiment of the present disclosure also provides a computer program product. The computer program product stores instructions that, when executed by a computer, cause the computer to implement the indoor navigation method described in the above method embodiments.

[0163] The above instructions may be program code. In specific implementation, the program code may be written in any combination of one or more programming languages.

[0164] The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages.

[0165] The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0166] In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through an Internet service provider via the Internet).

[0167] The embodiment of the present disclosure also provides a chip, including at least one processor and an interface;

[0168] An interface for providing program instructions or data to at least one processor;

[0169] At least one processor is configured to execute program instructions to implement the indoor navigation method described in the above method embodiments.

[0170] In some embodiments, the chip may further include a memory for storing program instructions and data, and the memory is located inside or outside the processor.

[0171] Those of ordinary skill in the art may understand that all or part of the steps to implement the above embodiments may be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which may be collectively referred to as "circuitry", "module", or "system" here.

[0172] Those skilled in the art will readily conceive of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein.

[0173] The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. An indoor navigation method, characterized in that: include: Obtaining a user's navigation demand, analyzing the navigation demand using a transformer-based multi-task learning system, determining a current location and a destination location that require navigation, wherein the current location and the destination location are located inside the same building, and the navigation demand includes one or more segments of voice data; Generate a navigation route based on the current location, the destination location and a topological structure library of the building, the navigation route including a plurality of points of interest between the current location and the destination location, the topological structure library being a database storing internal spatial layout information of the building; Through multiple rounds of dialogue between the virtual person and the user, the current location of the user is updated in real time; In the case that the current position does not deviate from the navigation route, in response to the update of the current position, the virtual person is controlled to issue a navigation voice instruction, wherein the navigation voice instruction is used to guide the user to go from the current position to the next point of interest of the navigation route.

2. The method according to claim 1, characterized in that The obtaining of the user's navigation needs, analyzing the navigation needs using a multi-task learning system based on a transformer architecture, and determining the current location and the destination location that require navigation include: The user's navigation needs are obtained through the dialogue between the virtual human and the user, and the position described by the user in the navigation needs in natural language is converted into a spatial coordinate system using a multi-task learning system based on the transformer architecture to obtain the analysis result; Through multiple rounds of dialogues between the virtual person and the user, it is confirmed whether the current location and the destination location in the analysis result are accurate. If inaccurate, the virtual person and the user are dialogued again and the multi-task learning system is used for analysis until the user confirms that the current location and the destination location in the analysis result are accurate.

3. The method according to claim 1, characterized in that The method of updating the current location of the user in real time through multiple rounds of dialogue between the virtual person and the user includes: The user's current estimated location is obtained in combination with the user's walking speed, and a conversation between the virtual person and the user is triggered before the user reaches the next point of interest to determine whether the estimated location is accurate, and the multi-task learning system is used to analyze the conversation and update the user's current location.

4. The method according to claim 1, characterized in that: The virtual person and the user have a conversation, including: Acquire voice data input by the user and a facial image of the user when performing voice input; Sending the voice data to a server, so that the server can obtain text semantics based on the voice data analysis, and obtain the user's voice emotion based on the intonation recognition of the voice data; Sending the facial image to a server, so that the server can recognize the user's visual emotion based on the facial image when the user performs voice input; In response to the first instruction from the server, the virtual person makes a voice reply; wherein, the reply voice of the virtual person's voice reply is generated based on the text semantics, the voice emotion and the visual emotion, and the intonation and facial expression of the virtual person's voice reply are adjusted based on the text semantics, the voice emotion and the visual emotion.

5. The method according to claim 1, characterized in that The generating of a navigation route based on the current position, the destination position and the topological structure library of the building comprises: On the virtual human interaction page, multiple routes to be selected are displayed, where the multiple routes to be selected are generated based on the current location, the destination location, and the topological structure library of the building; In response to a selection operation of a navigation route among the multiple routes to be selected, the navigation route is displayed.

6. The method according to claim 1, characterized in that After the current location of the user is updated in real time through multiple rounds of dialogues between the virtual person and the user, the method further includes: In the case where the current position deviates from the navigation route, generating a new navigation route based on the updated current position, the destination position and the topological structure library of the building; The virtual person is controlled to issue a navigation voice instruction, wherein the navigation voice instruction is used to guide the user to go from the updated current position to the next point of interest of the new navigation route.

7. The method according to claim 4, characterized in that The sending of the voice data to the server includes: Using ASR streaming technology to perform real-time encoding and segmentation processing on the voice data; The data after real-time encoding and segmentation processing is encapsulated through the RTC protocol and sent to the server.

8. The method according to claim 4, characterized in that When the virtual person makes a voice reply, the method further comprises: Use VAD technology to detect whether there is new voice activity or text input; If new voice activity or text input is detected, a semantic integrity assessment is initiated to obtain an assessment result; When the interruption response window time is reached, determining the user intention based on the voice activity and / or text input by the user within the interruption response window time, and determining whether an interruption is required based on the user intention; If interruption is required, determining an interruption point in the current context according to the evaluation result; The ongoing voice response of the virtual person is paused at the interruption point, and new voice activity or text input of the user is received.

9. The method according to claim 1, characterized in that: Before generating a navigation route based on the current location, the destination location and the topological structure library of the building, the method further includes: When the dynamic attributes associated with the destination location meet preset conditions, the virtual person communicates with the user to recommend a new destination location to the user; wherein the new destination location is determined based on the user's historical behavior; and the dynamic attributes include business hours and / or traffic volume.

10. An electronic device, characterized in that: include: A memory for storing instructions; A processor is used to call the instructions stored in the memory to implement the indoor navigation method as described in any one of claims 1-9.

11. A computer program product, characterized in that The computer program product stores instructions, and when the instructions are executed by a computer, the computer implements the indoor navigation method according to any one of claims 1 to 9.