Enhanced navigation termination method and device, computer equipment and storage medium
By performing keyframe screening of target areas and probability evaluation of visual language models, the exploration of low-value areas was terminated, and the problem of poor termination decisions in the existing navigation system was solved, and navigation efficiency and resource utilization were improved.
Patent Information
- Application Number
- CN202510652427.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-26
AI Technical Summary
The existing target navigation system has the problem of poor exploration and termination decisions in the medical and health field, resulting in waste of redundant exploration and computing resources, reducing navigation efficiency and user experience.
By exploring the target area, recording keyframes, filtering and sorting the keyframes with the highest contribution, using visual language models for macro perception, judging the probability of the target object's existence, and terminating exploration in low-value areas.
It improves exploration efficiency, reduces redundant exploration, optimizes computing resource allocation, and improves the overall performance and user experience of the navigation system.
Smart Images

Figure CN120544147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and safe travel technology, and in particular to a method, device, computer equipment and storage medium for terminating enhanced navigation. Background Art
[0002] In the field of healthcare (such as the elderly care field), existing target navigation systems have the following problems when assisting users in finding items: First, traditional navigation strategies usually require a complete search of the current area before moving to a new area, and cannot accurately evaluate the current exploration status. From the perspective of marginal utility, this strategy has the problem of low efficiency. Specifically, in the early stages of information collection, the input-output ratio is relatively high, but as the exploration deepens, the marginal value of the new information obtained from each operation gradually decreases. Secondly, existing research mainly relies on exploration maps constructed by visual perception to identify targetless areas. However, precision errors and model limitations make it difficult to fully annotate these areas on the map, resulting in the robot triggering repeated boundary settings due to small unknown areas, causing redundant exploration. This inefficient navigation method may increase user waiting time and reduce their user experience.
[0003] Furthermore, existing goal-oriented navigation systems in elderly care scenarios suffer from poor exploration termination decisions. Due to the lack of an effective exploration termination strategy, the system cannot accurately determine when to terminate exploration of the current area and move on to more valuable areas. This not only wastes time and computing resources, but also reduces the overall efficiency of navigation. For example, when a robot is helping an elderly person find their glasses, if it fails to promptly terminate the search in unhelpful areas, it may delay the user's search and even cause anxiety. Summary of the Invention
[0004] The purpose of the present invention is to provide a termination enhanced navigation method, apparatus, computer equipment and storage medium user, aiming to solve the problem of poor exploration termination decision-making in the target navigation system in the prior art.
[0005] In a first aspect, an embodiment of the present invention provides a method for terminating enhanced navigation, including:
[0006] exploring a target area, and recording a plurality of key frames of the target area during the exploration process;
[0007] Filtering and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution;
[0008] Performing macroscopic perception of the key frames in the key frame set through a visual language model to determine whether a target object exists in the target area, and directly outputting an exploration result if so; if not, evaluating the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability;
[0009] The probability result is outputted through the visual language model, and when the probability result is a third probability, the target area is marked as a low-value area, and the exploration of the target area is terminated.
[0010] In a second aspect, an embodiment of the present invention further provides a termination enhanced navigation device, comprising:
[0011] An exploration unit, configured to explore a target area and record a plurality of key frames of the target area during the exploration process;
[0012] a sorting unit, configured to screen and sort the plurality of key frames based on exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution;
[0013] An evaluation unit is configured to perform macroscopic perception of the key frames in the key frame set using a visual language model, determine whether a target object exists in the target area, and directly output an exploration result if so; if not, evaluate the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability;
[0014] The exploration termination unit is configured to output the probability result through the visual language model, and when the probability result is a third probability, mark the target area as a low-value area and terminate the exploration of the target area.
[0015] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for terminating enhanced navigation as described in the first aspect when executing the computer program.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the termination enhanced navigation method described in the first aspect.
[0017] The embodiments of the present invention provide a method, apparatus, computer equipment and storage medium for terminating enhanced navigation. The method can concentrate on processing key frames that contribute most to exploration by screening and sorting key frames, avoid processing a large amount of redundant information, and thus significantly improve exploration efficiency. With the help of the macro-perception ability and three-level probability evaluation mechanism of the visual language model, low-value areas can be identified in a timely manner and exploration can be terminated, avoiding unnecessary repeated searches in areas without target objects, thereby reducing redundant exploration and further improving navigation efficiency. At the same time, by promptly terminating the exploration of low-value areas, limited computing resources can be concentrated on target areas where the target object is more likely to be found, reducing the waste of computing resources, thereby optimizing the allocation of computing resources and improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 is a schematic diagram of an application environment of the method for terminating enhanced navigation provided in an embodiment of the present invention;
[0020] Figure 2 is a flow chart of a method for terminating enhanced navigation provided in an embodiment of the present invention;
[0021] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation method before step S102;
[0022] Figure 4 is a structural diagram of a termination enhanced navigation device provided in an embodiment of the present invention;
[0023] Figure 5 is a structural diagram of a computer device provided in an embodiment of the present invention;
[0024] Figure 6 2 is another structural diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0027] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0028] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] The termination enhanced navigation method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, wherein the client communicates with the program terminal through a network, the program terminal can explore the target area through the information of the target area provided by the client, and record multiple key frames of the target area during the exploration process; multiple key frames are screened and sorted based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution; the key frames in the key frame set are macro-perceived through a visual language model to determine whether there is a target object in the target area, and if so, the exploration result is directly output; otherwise, the probability of the target object existing in the target area is evaluated to obtain a probability result, wherein the probability result includes a first probability, a second probability and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability; the probability result is output through the visual language model, and when the probability result is the third probability, the target area is marked as a low-value area, and the exploration of the target area is terminated.
[0030] In the present invention, for complex insurance entities under the insurance business, especially scenarios involving pension insurance for the elderly, insurance companies can provide intelligent assistance services for the elderly to help them quickly find target objects (such as medicines, glasses, etc.) and improve the convenience of life. The target area is explored through the target area information provided by the user (such as room, living room, bedroom, etc.). During the exploration process, multiple key frames of the target area are recorded. These key frames may contain important visual information such as the corners of the room and the placement of furniture. By screening and sorting multiple key frames based on the exploration contribution, a key frame set is obtained, wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution. For example, in a pension environment, these key frames may be concentrated in areas where the elderly often move or areas where items are easily placed, such as bedside tables, desks, sofas, and other locations.
[0031] Next, the key frames in the key frame set are macro-perceived through the visual language model to determine whether there is a target object (such as medicines, glasses, remote controls, etc. required by the user) in the target area. If the visual language model determines that the target object exists, the exploration result is directly output to help the user quickly find the required items. If the target object is not found, the probability of the existence of the target object in the target area is further evaluated to obtain a probability result, where the probability result includes a high probability (first probability), an uncertain probability (second probability) and a low probability (third probability). If the probability result is a low probability (third probability), the target area is marked as a low-value area, and further exploration of the area is terminated, thereby avoiding wasting time and computing resources in areas without target objects. This navigation method can significantly improve the efficiency of finding target objects in the elderly care environment, reduce the waiting time of users, improve the convenience and comfort of users' lives, and also reduce the workload of caregivers.
[0032] For application scenarios in the medical field (such as elderly care), by screening and sorting key frames, it is possible to concentrate on processing the key frames that contribute the most to exploration, avoiding the processing of a large amount of redundant information, thereby significantly improving the efficiency of exploration. With the help of the macro-perception ability of the visual language model and the three-level probability evaluation mechanism, low-value areas can be identified in a timely manner and the exploration can be terminated, avoiding unnecessary repeated searches in areas without target objects, thereby reducing redundant exploration. Furthermore, by promptly terminating the exploration of low-value areas, limited computing resources can be concentrated in areas where it is more likely to find the target, thereby optimizing the allocation of computing resources and improving the overall performance of the system. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, and tablets. The program end can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0033] See also Figure 2, Figure 2 A flowchart of a method for terminating enhanced navigation provided by an embodiment of the present invention includes steps S101 to S104:
[0034] S101, exploring a target area, and recording a plurality of key frames of the target area during the exploration process;
[0035] In this embodiment, the target area can be any space to be searched, such as a room or multiple rooms in an indoor environment. The target area is explored, and multiple key frames of the target area are recorded during the exploration process. Key frames are image frames captured during the exploration process that contain important visual information and can reflect the characteristics and structure of the target area.
[0036] In a senior care setting, the target area could be the elderly person's living space, such as a bedroom, living room, or kitchen. A robot or smart device explores the target area and records keyframes. For example, keyframes in a bedroom might include images of locations like the bedside table, wardrobe, and window; keyframes in a living room might include images of locations like the sofa, coffee table, and TV stand; and keyframes in a kitchen might include images of cabinets, countertops, and refrigerators. These keyframes capture important visual information about the target area, providing foundational data for subsequent object detection and navigation.
[0037] S102, screening and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution;
[0038] In this embodiment, the multiple keyframes recorded during the exploration process are further processed. Specifically, each keyframe is screened and ranked based on its contribution to target area exploration. Exploration contribution refers to the degree to which a keyframe contributes to target area information coverage and target detection during the exploration process. This evaluation identifies the most valuable keyframes and groups them into a keyframe set. The number of keyframes in this set is pre-set, typically an optimal value determined based on task requirements and computational resource constraints.
[0039] Specifically, in a senior living environment, let's assume the target area is a senior's bedroom. During exploration, multiple keyframes are recorded, potentially including images of locations such as a bedside table, wardrobe, and window. The exploration contribution of these keyframes is evaluated: If the bedside table is where the senior often places items (such as glasses or medicine), then this keyframe's exploration contribution will be high, as it significantly helps quickly find the target object. If the wardrobe is large and the items are complexly arranged, but the target object is unlikely to be there, its contribution may be lower. A keyframe for a window: If no items are placed near the window, its contribution will also be lower. The keyframes are then sorted based on these contributions, ultimately selecting a predetermined number of keyframes with the highest contributions to form a keyframe set. This approach allows for focused processing of the most valuable keyframes, avoiding the processing of redundant information. This improves navigation efficiency, reduces waiting time for seniors, and enhances their user experience.
[0040] Among them, Figure 3 As shown, in step S102, before the step of screening and sorting the multiple key frames based on the exploration contribution to obtain a key frame set, the following steps S201 to S203 are included:
[0041] S201, obtaining the current area exploration rate of the target area, and calculating the exploration rate threshold of the target area;
[0042] S202: Determine whether the current area exploration rate of the target area exceeds the exploration rate threshold;
[0043] S203. If yes, trigger the macro-perception function of the visual language model and perform the key frame screening and sorting steps; if no, do not trigger the macro-perception function of the visual language model and do not perform the key frame screening and sorting steps.
[0044] In step S201, the region exploration rate refers to the proportion of the target region that has been explored, and is generally used to measure the progress of the current exploration. At the same time, a target region exploration rate threshold is also calculated. The exploration rate threshold is a preset value used to determine whether the current exploration has achieved sufficient coverage, thereby determining whether to trigger the macro-perception function of the visual language model (VLM).
[0045] Furthermore, in step S202, it is determined whether the current region exploration rate exceeds the calculated exploration rate threshold. If the current region exploration rate exceeds the threshold, it indicates that the target region has been sufficiently explored and the macro-perception function of the visual language model can be triggered. If it does not exceed the threshold, it indicates that the exploration is not sufficient and the macro-perception function of the visual language model is not triggered for the time being.
[0046] In step S203, if the current region exploration rate exceeds the exploration rate threshold, the macro-perception function of the visual language model is triggered, and the keyframe screening and sorting steps are executed. In other words, the current exploration has sufficiently covered the target area, and the keyframes can be further processed using the visual language model to determine whether the target object exists or to assess the probability of its existence. If the current region exploration rate does not exceed the threshold, the macro-perception function of the visual language model is not triggered, and the keyframe screening and sorting steps are not executed. Instead, exploration continues.
[0047] Wherein, in step S201, the step of calculating the exploration rate threshold of the target area includes S301:
[0048] S301. Calculate the exploration rate threshold of the target area according to the following formula:
[0049] τ t =τ0-β·Δt+γ·Complexity(R t )
[0050] Among them, τ t represents the exploration rate threshold at time t, τ0 represents the initial threshold, β represents the time attenuation coefficient, Δt represents the time step passed, γ represents the complexity adjustment coefficient, Complexity(R t ) represents the target area R t complexity.
[0051] In this embodiment, a dynamic exploration rate threshold adjustment mechanism is introduced, enabling the system to adaptively adjust the threshold for triggering the VLM macro-perception function based on environmental complexity and task difficulty. The core of this dynamic adjustment mechanism is that it does not use a fixed threshold, but rather flexibly adjusts the threshold based on the specific conditions of the current exploration environment to ensure that the VLM macro-perception function can be efficiently triggered in different scenarios.
[0052] Furthermore, in step S301, calculating the complexity of the target area includes step S401:
[0053] S401. Calculate the complexity of the target area according to the following formula:
[0054]
[0055] Among them, w1, w2 and w3 are weight coefficients, Indicates the relative size of the target area, max i |R i | indicates the size of the largest target area among multiple target areas. represents the shape complexity of the target area, Indicates the density of objects in the target area.
[0056] Specifically, in a senior care setting, the target area might be the elderly person's living space, such as the bedroom, living room, or kitchen. The complexity of these areas can vary depending on the room layout, furniture placement, and number of items. By dynamically adjusting the exploration rate threshold, the system can better adapt to different environments and task requirements.
[0057] For example, the target area is a simple bedroom with a small area, neatly arranged furniture, and low object density. Then the complexity of the target area (R t ) is lower, and then a lower exploration rate threshold τ is calculated according to the formula t , it is possible to reach the conditions for triggering VLM macroscopic perception in a relatively short time, thereby quickly judging whether the target object exists. For another example, the target area is a complex living room with a large area, complex furniture arrangement, and high object density. Then the complexity of the target area (Complexity t ) is higher, then a higher exploration rate threshold τ is calculated according to the formula t , more exploration time is required to reach the conditions for triggering VLM macroscopic perception, thereby ensuring a more comprehensive assessment of the target area in complex environments.
[0058] In summary, the dynamic adjustment mechanism of the exploration rate threshold allows us to adaptively adjust the exploration rate threshold based on the complexity of the target area and the difficulty of the task, ensuring that the VLM macro-perception function can be efficiently triggered in different environments. Furthermore, by properly adjusting the exploration rate threshold, we can reduce unnecessary exploration time while ensuring exploration quality, thereby improving navigation efficiency.
[0059] In the step S102, the step of screening and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set includes the following steps S501 to S502:
[0060] S501. Filter the plurality of key frames according to the following formula to obtain a filtered key frame set:
[0061]
[0062] Among them, f i represents the i-th key frame, R(f i ) represents the field of view of the i-th key frame, R t represents the target area, K represents K key frames, represents the empty set;
[0063] S502: Calculate the exploration contribution of each key frame in the filtered key frame set according to the following formula:
[0064]
[0065] Among them, C(f i ) represents the i-th key frame f i The exploration contribution, V(f i ) represents the i-th key frame f i visible area.
[0066] In this embodiment, multiple key frames are screened, and the screening condition is that the field of view of the key frame must be within the target area, ensuring that only key frames that actually cover the target area are retained, thereby improving the efficiency of subsequent processing.
[0067] Furthermore, the exploration contribution is calculated by the ratio of the visible area of the keyframe to the target area. This ratio reflects the coverage of the target area by each keyframe, thus evaluating the value of each keyframe in the exploration process.
[0068] Furthermore, in step S102, the step of screening and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set further includes the following step S601:
[0069] S601: Select the N key frames with the highest contribution according to the following formula and combine them into a final key frame set:
[0070] S' f ={f i |f i ∈S f , sort C(f i ))≤N}
[0071] Among them, S' f Represents the final set of keyframes.
[0072] In this embodiment, after calculating the exploration contribution of each key frame in the key frame set filtered by the above calculation, the top N key frames with the highest contribution are selected and combined into a final key frame set to ensure that only the key frames that contribute the most to the exploration of the target area are used in subsequent processing, which can reduce redundant information and improve the efficiency of subsequent processing.
[0073] S103, performing macroscopic perception of the key frames in the key frame set using a visual language model to determine whether a target object exists in the target area, and if so, directly outputting an exploration result; if not, evaluating the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability;
[0074] In this embodiment, the visual language model first determines whether there is a target object in the target area. If the target object exists, the exploration results, such as the location information of the target object, are directly output. If the visual language model (VLM) cannot directly confirm the existence of the target object, the probability of the target object existing in the target area is evaluated, and the following three probability results are obtained: First probability (high probability): The probability of the target object existing is very high. Second probability (uncertain probability): The probability of the target object existing is uncertain and further exploration is required. Third probability (low probability): The probability of the target object existing is low.
[0075] Furthermore, based on these probability results, the following decisions can be made: If the target area is evaluated as having a high probability (first probability), the system can prioritize exploration of that area. If the target area is evaluated as having an uncertain probability (second probability), the system can mark that area as requiring further exploration. If the target area is evaluated as having a low probability (third probability), the system can mark that area as low-value and terminate further exploration of that area.
[0076] Wherein, in step S103, evaluating the probability of the target object existing in the target area to obtain a probability result includes the following step S701:
[0077] S701: Model the output of the visual language model using a Bayesian probability framework according to the following formula:
[0078]
[0079] Among them, P(O|I,E) represents the posterior probability of the existence of the target object O given the image input I and the information E of the target area, P(I|O,E) represents the likelihood probability of observing the image I when the target object O exists and the information E of the target area is given, P(O|E) represents the prior probability of the existence of the target object O given the information E of the target area, and P(I|E) represents the edge probability of observing the image I given the information E of the target area.
[0080] In this embodiment, the Bayesian probability framework is a statistical method used to update the probability of a hypothesis (such as the existence of a target object) given new evidence (such as an image input). Specifically, the system calculates the posterior probability P(O|I,E) of the target object's existence, which is calculated based on the likelihood probability, prior probability, and marginal probability.
[0081] Specifically, in a senior care environment, assume the target area is a senior's bedroom, and the target object is medication they need. In step S103, a macroscopic perception of the keyframe set is performed using the VLM, and the VLM output is obtained. In step S701, the VLM output is modeled using a Bayesian probabilistic framework to calculate the posterior probability P(O|I,E) of the medication's presence.
[0082] For example, if the VLM output indicates a high probability of observing medication in the bedside table area, the likelihood probability P(I|P,E) will be relatively high. Furthermore, based on prior knowledge (e.g., elderly people often place medication on their bedside tables), the prior probability P(O|E) will also be relatively high. Therefore, the posterior probability P(O|I,E) of the medication's presence will be relatively high, indicating a high probability of its presence.
[0083] S104: Output the probability result through the visual language model, and when the probability result is a third probability, mark the target area as a low-value area, and terminate the exploration of the target area.
[0084] In this embodiment, if the probability result is the third probability (low probability), the target area is marked as a low-value area and further exploration of the area is terminated. This mechanism can effectively reduce redundant exploration and improve the overall efficiency of the system.
[0085] Wherein, in the step S104, the step of marking the target area as a low-value area includes S801:
[0086] S801: Mark the target area as a low-value area according to the following formula:
[0087]
[0088] Among them, S p represents the exploration priority score of the target area, α represents the decay coefficient less than 1, S base Represents the base exploration priority score.
[0089] In this embodiment, the exploration priority score is an indicator used to evaluate the exploration value of a target area. When a target area is marked as low-value, its exploration priority score is reduced according to the decay coefficient α. This allows the system to prioritize areas with higher exploration priority scores in subsequent explorations, thereby improving navigation efficiency.
[0090] It can be seen that in the above scheme, for complex elderly care entities such as elderly care businesses, by exploring the target area, screening and sorting multiple key frames, evaluating the probability of the target object existing in the target area, and finally terminating the exploration of low-value areas, the problem of poor exploration termination decision-making in the target navigation system in the prior art is solved, and the processing of a large amount of redundant information is avoided, thereby significantly improving the exploration efficiency. At the same time, it can promptly identify low-value areas and terminate the exploration, avoiding unnecessary repeated searches in areas without target objects, thereby reducing redundant exploration and further improving navigation efficiency. It can concentrate limited computing resources on target areas where there is a greater hope of finding target objects, reducing the waste of computing resources, thereby optimizing the allocation of computing resources and improving the overall performance of the system.
[0091] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0092] The embodiment of the present invention further provides a device for terminating enhanced navigation, which corresponds to the method for terminating enhanced navigation in the above embodiment. Figure 4 As shown, the termination enhanced navigation device 900 includes an exploration unit 901, a sorting unit 902, an evaluation unit 903, and a termination exploration unit 904. The functional units are described in detail as follows:
[0093] An exploration unit 901 is configured to explore a target area and record a plurality of key frames of the target area during the exploration process;
[0094] A sorting unit 902 is configured to screen and sort the plurality of key frames based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution;
[0095] An evaluation unit 903 is configured to perform macroscopic perception of the key frames in the key frame set using a visual language model to determine whether a target object exists in the target area. If so, the evaluation unit 903 directly outputs an exploration result. If not, the evaluation unit 903 evaluates the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability. The second probability is an uncertain probability, and the first probability is greater than the third probability.
[0096] The exploration termination unit 904 is configured to output the probability result through the visual language model, and when the probability result is the third probability, mark the target area as a low-value area and terminate the exploration of the target area.
[0097] The present invention provides a termination-enhanced navigation device that, by screening and sorting key frames, can focus on processing key frames that contribute most to exploration, avoiding the processing of large amounts of redundant information, thereby significantly improving exploration efficiency. By leveraging the macroscopic perception capabilities and three-level probability evaluation mechanism of the visual language model, low-value areas can be promptly identified and exploration terminated, avoiding unnecessary repeated searches in areas without target objects, thereby reducing redundant exploration and further improving navigation efficiency. At the same time, by promptly terminating the exploration of low-value areas, limited computing resources can be concentrated on target areas where the target object is more likely to be found, reducing the waste of computing resources, thereby optimizing the allocation of computing resources and improving the overall performance of the system.
[0098] The specific limitations of the enhanced navigation termination device can be found in the limitations of the enhanced navigation termination method described above and will not be further elaborated here. Each module in the enhanced navigation termination device described above may be implemented in whole or in part via software, hardware, or a combination thereof. Each of the modules may be embedded in or independent of a processor within a computer device in hardware form, or may be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0099] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements a function or step on the server side of a method for terminating enhanced navigation.
[0100] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a function or step on the client side of a method for terminating enhanced navigation.
[0101] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0102] exploring a target area, and recording a plurality of key frames of the target area during the exploration process;
[0103] Filtering and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution;
[0104] Performing macroscopic perception of the key frames in the key frame set through a visual language model to determine whether a target object exists in the target area, and directly outputting an exploration result if so; if not, evaluating the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability;
[0105] The probability result is outputted through the visual language model, and when the probability result is a third probability, the target area is marked as a low-value area, and the exploration of the target area is terminated.
[0106] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0107] exploring a target area, and recording a plurality of key frames of the target area during the exploration process;
[0108] Filtering and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution;
[0109] Performing macroscopic perception of the key frames in the key frame set through a visual language model to determine whether a target object exists in the target area, and directly outputting an exploration result if so; if not, evaluating the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability;
[0110] The probability result is outputted through the visual language model, and when the probability result is a third probability, the target area is marked as a low-value area, and the exploration of the target area is terminated.
[0111] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0112] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0113] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0114] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for terminating enhanced navigation, characterized in that: The following steps are involved: exploring a target area, and recording a plurality of key frames of the target area during the exploration process; Filtering and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution; Performing macroscopic perception of the key frames in the key frame set through a visual language model to determine whether a target object exists in the target area, and directly outputting an exploration result if so; if not, evaluating the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability; The probability result is outputted through the visual language model, and when the probability result is a third probability, the target area is marked as a low-value area, and the exploration of the target area is terminated.
2. The termination enhanced navigation method according to claim 1, characterized in that: Before the plurality of key frames are screened and sorted based on the exploration contribution to obtain a key frame set, the following steps are included: Obtaining a current area exploration rate of the target area and calculating an exploration rate threshold of the target area; Determining whether a current area exploration rate of the target area exceeds the exploration rate threshold; If so, the macro-perception function of the visual language model is triggered, and the key frame screening and sorting steps are performed; if not, the macro-perception function of the visual language model is not triggered, and the key frame screening and sorting steps are not performed.
3. The termination enhanced navigation method according to claim 2, characterized in that: The plurality of key frames are screened and sorted based on the exploration contribution to obtain a key frame set, including: The plurality of key frames are filtered according to the following formula to obtain a filtered key frame set: Among them, f i represents the i-th key frame, R(f i ) represents the field of view of the i-th key frame, R t represents the target area, K represents K key frames, represents the empty set; The exploration contribution of each key frame in the filtered key frame set is calculated according to the following formula: Among them, C(f i ) represents the i-th key frame f i The exploration contribution, V(f i ) represents the i-th key frame f i visible area.
4. The termination enhanced navigation method according to claim 3, characterized in that: The step of screening and sorting the plurality of key frames based on the exploration contribution to obtain a key frame set further includes: Select the N keyframes with the highest contribution according to the following formula and combine them into the final keyframe set: S' f ={f i |f i ∈S f , sort(C(f i ))≤N} Among them, S' f Represents the final set of keyframes.
5. The termination enhanced navigation method according to claim 2, characterized in that: The step of calculating the exploration rate threshold of the target area includes: The exploration rate threshold of the target area is calculated according to the following formula: t t =τ0-β·Δt+γ·Complexity(R t ) Among them, τ t represents the exploration rate threshold at time t, τ0 represents the initial threshold, β represents the time attenuation coefficient, Δt represents the time step passed, γ represents the complexity adjustment coefficient, Complexity(R t ) represents the target area R t complexity.
6. The termination enhanced navigation method according to claim 5, characterized in that: The complexity of the target area is calculated according to the following formula: Among them, w1, w2 and w3 are weight coefficients, Indicates the relative size of the target area, max i |R i | indicates the size of the largest target area among multiple target areas. represents the shape complexity of the target area, Indicates the density of objects in the target area.
7. The termination enhanced navigation method according to claim 1, characterized in that: The evaluating the probability of the target object existing in the target area to obtain a probability result includes: The output of the visual language model is modeled using a Bayesian probabilistic framework according to the following formula: Among them, P(O|I,E) represents the posterior probability of the existence of the target object O given the image input I and the information E of the target area, P(I|O,E) represents the likelihood probability of observing the image I when the target object O exists and the information E of the target area is given, P(O|E) represents the prior probability of the existence of the target object O given the information E of the target area, and P(I|E) represents the edge probability of observing the image I given the information E of the target area.
8. A termination enhanced navigation device, characterized in that: include: An exploration unit, configured to explore a target area and record a plurality of key frames of the target area during the exploration process; a sorting unit, configured to screen and sort the plurality of key frames based on exploration contribution to obtain a key frame set; wherein the key frame set includes a predetermined number of key frames with the highest exploration contribution; An evaluation unit is configured to perform macroscopic perception of the key frames in the key frame set using a visual language model, determine whether a target object exists in the target area, and directly output an exploration result if so; if not, evaluate the probability of the target object existing in the target area to obtain a probability result, wherein the probability result includes a first probability, a second probability, and a third probability; wherein the second probability is an uncertain probability, and the first probability is greater than the third probability; The exploration termination unit is configured to output the probability result through the visual language model, and when the probability result is a third probability, mark the target area as a low-value area and terminate the exploration of the target area.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the termination enhanced navigation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the termination enhanced navigation method according to any one of claims 1 to 7.