Method, device, medium and equipment based on open 3D (three-dimensional) large-world no-graph navigation
By using a fusion network to predict the actions of intelligent agents and adjust network parameters in open 3D large-world games, the problems of poor navigation and obstacle avoidance capabilities and low task completion efficiency in existing technologies are solved, achieving accurate navigation and efficient task completion.
Patent Information
- Application Number
- CN202510711877.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing open 3D large-world game navigation technology relies on static maps or traditional reinforcement learning algorithms, which cannot adapt to dynamic obstacles in real time, have poor navigation and obstacle avoidance capabilities, cannot accurately navigate intelligent agents, and have low task completion efficiency.
By fusing the image features of the network input game interface and the pose vector features of the agent, the next action of the agent is predicted, and the reward value is determined by analyzing the actual position, and the network parameters are adjusted to optimize the navigation model.
It achieves precise navigation of intelligent bodies in an open 3D world, improves task completion efficiency and user gaming experience, and reduces dependence on static maps.
Smart Images

Figure CN120617952A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of game navigation technology, and more specifically, to a method, apparatus, medium, and device for mapless navigation in an open 3D world. Background Art
[0002] Open 3D world games are electronic games that offer vast game worlds where players can freely explore, interact, and complete tasks. The core feature of these games is providing players with a vast, freely explorable environment where they can freely choose their goals and paths, interact with other characters, and engage in various activities. Currently, navigation in open 3D world games primarily utilizes two approaches: static map navigation based on relevant pathfinding algorithms, which requires pre-loading a complete, pre-built scene map; and basic reinforcement learning (RL) algorithms trained using the Deep Q-Network (DQN) or the original Proximal Policy Optimization (PPO) algorithm, which takes RGB images as input and relies on recognizing in-game pathfinding cues (such as minimap targets and route prompts). However, both approaches rely heavily on a static environment, lack real-time adaptation to dynamic obstacles, and exhibit poor obstacle avoidance capabilities. This inability to accurately guide the agent in a task results in low task completion efficiency.
[0003] Therefore, how to provide a technical solution for a method of mapless navigation in an open 3D world with high accuracy has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The purpose of some embodiments of the present application is to provide a method, apparatus, medium, and equipment for mapless navigation in an open 3D world. The technical solutions of the embodiments of the present application can improve the accuracy of mapless navigation in open 3D world games, allowing intelligent agents to find their way correctly, thereby improving task completion efficiency and user gaming experience.
[0005] In a first aspect, some embodiments of the present application provide a method for mapless navigation based on an open 3D world, comprising: inputting the image features of a game interface currently displayed on the screen and the pose vector features of an intelligent agent into a fusion network to predict the next action of the intelligent agent; wherein, the image features include: object information and the relative distance between objects; the next action includes: the action type and moving direction of the intelligent agent; by analyzing the actual position of the intelligent agent after performing the next action, a reward value is determined; wherein, the reward value includes a positive reward value and a negative reward value; the positive reward value represents that the intelligent agent is gradually approaching the task target point, and the negative reward value represents that the intelligent agent is gradually moving away from the task target point or gradually approaching the target obstacle; based on the reward value, the network parameters of the fusion network are adjusted to obtain a target fusion model; wherein, the target fusion model is used to provide action and direction navigation for the intelligent agent when the intelligent agent performs a game task.
[0006] Some embodiments of the present application use a fusion model to predict the next action of an intelligent agent through the image features of the current screen and the posture feature vector of the intelligent agent; determine the reward value for the next action by analyzing the actual position of the intelligent agent after performing the next action, and then adjust the network parameters of the fusion network according to the reward value to obtain a target fusion model, thereby obtaining a target fusion model for precise navigation of the intelligent agent, and improving the accuracy of map-free navigation in open 3D large-world games, so that the intelligent agent can find the path correctly, improving task completion efficiency and user gaming experience.
[0007] In some embodiments, before inputting the image features of the game interface displayed on the current screen and the posture vector features of the intelligent body into the fusion network, the method further includes: obtaining image information of the game interface displayed on the current screen, and obtaining the posture information of the intelligent body in the current screen; analyzing the image information to determine the image features in the image information; and generating the posture vector features corresponding to the posture information.
[0008] Some embodiments of the present application generate image features by analyzing the image information currently displayed on the screen, convert the posture information of the intelligent body into posture vector features, and provide data support for the prediction of the intelligent body's movements.
[0009] In some embodiments, the image features include: an instance segmentation map and a relative depth map; the analyzing the image information to determine the image features in the image information includes: inputting the image information into a pre-trained instance segmentation network to obtain the instance segmentation map; wherein the instance segmentation map represents the object information of various objects contained in the image information; inputting the image information into a pre-trained monocular depth network to obtain the relative depth map; wherein the relative depth map represents the relative distances between various objects.
[0010] Some embodiments of the present application input image information into different pre-trained instance segmentation networks and monocular depth networks respectively to obtain instance segmentation maps and relative depth maps, thereby achieving effective segmentation and extraction of image information.
[0011] In some embodiments, the reward value is determined by analyzing the actual position of the agent after performing the next action, including: when it is determined that the relative distance between the actual position and the task target point is less than a distance threshold, triggering the generation of the positive reward value.
[0012] Some embodiments of the present application can ensure the accuracy of the moving direction of the agent by determining the distance relationship between the actual position of the agent and the task target point and generating a reward value.
[0013] In some embodiments, when there is a height difference between the task target point and the current position of the intelligent agent, the actual position includes a horizontal position and a vertical height from the ground; the reward value is determined by analyzing the actual position of the intelligent agent after performing the next action, including: if the horizontal distance between the horizontal position of the intelligent agent and the task target point is not less than a distance threshold, the reward value is calculated according to a preset rule; if the horizontal distance between the horizontal position and the task target point is less than a distance threshold and the vertical distance between the intelligent agent and the task target point is greater than a height threshold, the reward value is generated in the vertical direction; if the horizontal distance between the horizontal position and the task target point is less than a distance threshold and the vertical height from the task target point is not greater than a height threshold, the reward value is generated in the horizontal direction.
[0014] Some embodiments of the present application generate different reward values by analyzing the distance relationship between the agent and the task target point in the horizontal and vertical directions, thereby ensuring that the agent can move accurately in scenes with height differences.
[0015] In some embodiments, the reward value generated in the vertical direction is within a set reward range.
[0016] Some embodiments of the present application avoid reward shocks by setting a reward range.
[0017] In some embodiments, the reward value is determined by analyzing the actual position of the agent after performing the next action, including: confirming that the distance between the actual position and the target obstacle is less than the distance between the current position of the agent and the target obstacle, then generating the negative reward value; the negative reward value represents that the agent is gradually approaching the target obstacle.
[0018] Some embodiments of the present application can improve the obstacle avoidance capability of the intelligent agent by analyzing the actual position of the intelligent agent and the distance between obstacles.
[0019] In a second aspect, some embodiments of the present application provide a device for mapless navigation based on an open 3D world, including: a fusion module, used to input the image features of the game interface currently displayed on the screen and the posture vector features of the agent into a fusion network to predict the next action of the agent; wherein, the image features include: object information and the relative distance between objects; the next action includes: the action type and movement direction of the agent; a reward module, used to determine the reward value by analyzing the actual position of the agent after performing the next action; wherein, the reward value includes a positive reward value and a negative reward value; the positive reward value represents that the agent is gradually approaching the task target point, and the negative reward value represents that the agent is gradually moving away from the task target point or gradually approaching the target obstacle; an optimization module, used to adjust the network parameters of the fusion network based on the reward value to obtain a target fusion model; wherein, the target fusion model is used to provide action and direction navigation for the agent when the agent performs a game task.
[0020] In a third aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.
[0021] In a fourth aspect, some embodiments of the present application provide an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor can implement a method as described in any embodiment of the first aspect when executing the program.
[0022] In a fifth aspect, some embodiments of the present application provide a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the following is a brief introduction to the drawings required for use in some embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0024] Figure 1 A diagram of a system for open 3D world mapless navigation provided in some embodiments of the present application;
[0025] Figure 2 One of the flow charts of the method for mapless navigation based on an open 3D world provided in some embodiments of the present application;
[0026] Figure 3 Flowchart 2 of a method for mapless navigation based on an open 3D world provided in some embodiments of the present application;
[0027] Figure 4 A block diagram of the apparatus for open 3D world mapless navigation provided in some embodiments of the present application;
[0028] Figure 5 A schematic diagram of an electronic device is provided for some embodiments of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.
[0030] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0031] In the related technologies, navigation in open 3D large-world games mostly relies on pre-generated NavMesh or traditional reinforcement learning algorithms (such as DQN, etc.). However, the existing technologies are dependent on static environments, that is, NavMesh needs to be pre-built and cannot adapt to dynamic obstacles in real time, and static map navigation requires a detailed and complete global map. Traditional methods lack detailed maps of different indoor planes, and traditional reinforcement learning algorithm solutions lack a reward mechanism for the vertical direction (z-axis), resulting in the inability to effectively identify stairs or elevator paths, and a high failure rate for cross-floor tasks. Traditional algorithms cannot handle static and dynamic obstacles well, resulting in AI easily getting stuck in dynamic environments, poor adjustment capabilities when encountering emergencies (such as being hit by a car in an urban scene), poor obstacle avoidance capabilities, and easy loss of direction, leading to task failure. Traditional algorithms rely on RBG image input, which causes the network to be overloaded, consumes too much resources, has poor obstacle recognition capabilities, is difficult to learn effective strategies, and has low utilization of visual information.
[0032] In view of this, some embodiments of the present application provide a method for mapless navigation based on an open 3D world. The method predicts the next action of the agent by inputting the image features of the current game interface and the pose vector features of the agent into the fusion network; and determines the reward value by analyzing the actual position of the agent after performing the next action. The reward value can evaluate the accuracy of the next action generation, and use this as a basis to optimize the fusion network to obtain a target fusion model. Some embodiments of the present application can analyze the image features of the current screen in the open 3D world mapless game scene in real time through the target fusion model, which can realize real-time recognition of the current game scene, improve the flexible adjustment ability of the agent, realize accurate navigation of the agent, improve task completion efficiency and user game experience, and reduce dependence on static maps.
[0033] The following is combined with Figure 1 The overall structure of the system based on open 3D world mapless navigation provided by some embodiments of the present application is exemplified.
[0034] like Figure 1As shown, some embodiments of the present application provide a system diagram based on open 3D world mapless navigation, and the system based on open 3D world mapless navigation may include: a terminal 100 and a server 200. The terminal 100 can send a screenshot of the game interface currently displayed on the screen and the current posture information of the agent to the server 200. The server 200 can analyze the screenshot to obtain image features; convert the posture information into posture vector features; then input the image features and posture vector features into the fusion network to be optimized, and output the next action of the agent; after the agent performs the next action, based on the actual position of the agent, determine the reward value, optimize the fusion network through the reward value, and obtain a target fusion model. Subsequently, the target fusion model can provide the agent with accurate direction navigation and action guidance in the open 3D world mapless game scene, and efficiently complete the game task.
[0035] In some embodiments of the present application, the terminal 100 may be a mobile terminal or a non-portable computer terminal, which is not specifically limited in the embodiments of the present application.
[0036] The following is combined with Figure 2 The implementation process of the open 3D world mapless navigation performed by the server 200 provided in some embodiments of the present application is exemplified.
[0037] Please see the attached Figure 2 , Figure 2 A flowchart of a method for imageless navigation based on an open 3D world is provided for some embodiments of the present application. The method for imageless navigation based on an open 3D world includes:
[0038] S210, inputting the image features of the game interface currently displayed on the screen and the posture vector features of the intelligent agent into the fusion network to predict the next action of the intelligent agent; wherein, the image features include: object information and the relative distance between objects; the next action includes: the action type and movement direction of the intelligent agent.
[0039] For example, in some embodiments of the present application, the image features of the currently displayed game interface and the pose vector features of the agent are cascaded, and feature fusion and dimensionality compression are performed through a fusion network composed of 256-dimensional fully connected layers to output the next action of the agent. Among them, the image features may include information such as different objects (such as trees, houses, vehicles, etc.) in the game interface on the current screen and the relative distances between objects in space. The content contained in the next action includes the action type (such as walking, running, squatting, jumping, etc.) and the direction of movement (such as front, back, left, right, left front, left rear, right front, right rear, etc.). The pose vector features may include 11-dimensional data such as the agent's own position, target position, orientation angle, distance, obstacle signs, etc. The pose vector feature structure can be simplified (such as eliminating the orientation angle) or expanded (adding speed, acceleration, etc.) according to the actual platform or task. The image features and the content contained in the next action can be flexibly set, and the embodiments of the present application are not limited to this. The fusion network can be an LSTM (Long-Short Term Memory) fusion model.
[0040] In some embodiments of the present application, before executing S210, the method based on open 3D large world mapless navigation includes: S201, obtaining image information of the game interface displayed on the current screen, and obtaining the posture information of the intelligent body in the current screen; analyzing the image information to determine the image features in the image information; generating the posture vector features corresponding to the posture information.
[0041] For example, in some embodiments of the present application, by taking a screenshot of the content displayed on the current screen, the image information displayed on the screen is obtained. The image information is analyzed and processed using a multimodal parallel processing architecture to obtain image features. The intelligent body's own coordinates, task target point coordinates, distance, azimuth, whether there are obstacles in front and other information constitute posture information, and then the posture information is converted into posture vector features through a fully connected layer network (which may be called a vector branch); wherein, the posture information may be normalized; the fully connected layer network maps the multidimensional state vector (i.e., posture information) to a high-dimensional feature space, and uses the ReLU activation function to enhance the nonlinear expression capability. The multimodal parallel processing architecture (which may be called the image processing branch) may include a trained instance segmentation network and a monocular deep network. It should be noted that the image processing branch and the vector branch may be a Transformer structure. The convolutional structure of the image processing branch (Conv2d 8→16→32→64) can be replaced with MobileNet or ResNet variants to adapt to different computing power platforms. The vector branch can be composed of different MLP depths or structures as long as the embedding mapping and feature expression of the state vector can be completed.
[0042] In some embodiments of the present application, image features include: an instance segmentation map and a relative depth map; 201 may include: inputting the image information into a pre-trained instance segmentation network to obtain the instance segmentation map; wherein the instance segmentation map represents the object information of various objects contained in the image information; inputting the image information into a pre-trained monocular depth network to obtain the relative depth map; wherein the relative depth map represents the relative distance between various objects.
[0043] For example, in some embodiments of the present application, the instance segmentation network can adopt the structure of YOLO11 real-time reasoning + channel separation. The image information obtained by the screenshot is respectively input into the trained instance segmentation network and the monocular depth estimation network (i.e., the monocular depth network) to realize the object segmentation and relative distance determination in the image information. For example, the instance segmentation network can process the obstacle information, feasible area information and other relevant information contained in the image information. For example, the box is an obstacle and an area that cannot be passed through, and the road is an area that can be walked, so that the intelligent agent can make decisions, avoid obstacles, and take a feasible path. The network in the multimodal parallel processing architecture can adopt a four-level progressive convolution layer structure, and each level of convolution layer is followed by a LayerNorm normalization layer and a ReLU activation function. The convolution step size is designed to be 2, and the size of the feature map output by each level of convolution layer is gradually compressed, and finally expanded into a multi-dimensional feature vector (as a specific example of image features) through the Flatten layer.
[0044] Among them, the instance segmentation network can also adopt a lightweight solution of YOLO+category mask post-processing; the monocular depth estimation network can be replaced by binocular vision, lidar depth map (such as in actual robot deployment) or other depth reconstruction technologies.
[0045] S220, determining a reward value by analyzing the actual position of the agent after performing the next action; wherein the reward value includes a positive reward value and a negative reward value; the positive reward value represents that the agent is gradually approaching the task target point, and the negative reward value represents that the agent is gradually moving away from the task target point or gradually approaching the target obstacle.
[0046] For example, in some embodiments of the present application, after the next action of the agent is output through the fusion network, the actual position of the agent after executing the next action can be evaluated to determine the effectiveness of the next action and give a corresponding reward value. That is, the subsequent series of rewards obtained are judged according to the next action currently selected, forming a similar total reward to provide feedback to the fusion network, which can be understood as evaluating the rationality of this action decision or this path through the reward value. For example, if the reward value is a positive reward value, it indicates that the action decision is correct and is gradually approaching the task target point; if the reward value is a negative reward value, it indicates that the action decision is wrong, and it may be gradually moving away from the task target point or approaching an obstacle, which hinders the completion of the game task.
[0047] It should be noted that when analyzing image features and posture vector features, the fusion network can calculate the action space probability values of multiple actions that the intelligent agent can perform, select the action corresponding to the maximum value of the action space probability value (i.e., the next action) as the output, and realize action space decision-making.
[0048] In some embodiments of the present application, the types of reward values may include: target point achievement reward, dynamic distance reward, and obstacle reward, etc. Different types of rewards may correspond to different reward functions. The three rewards are set with a reward priority order, for example, target point achievement reward > obstacle reward > dynamic distance reward. In actual applications, at least one of the three rewards can be selected to optimize the fusion model. In this case, a multi-stage reward function can be set, which is not specifically limited in the embodiments of the present application.
[0049] In some embodiments of the present application, S220 may include: triggering the generation of the positive reward value when determining that the relative distance between the actual position and the task target point is less than a distance threshold.
[0050] For example, in some embodiments of the present application, for the reward for reaching the target point, the reward value is determined by calculating the Euclidean distance between the actual position of the agent and the task target point (as a specific example of relative distance) and comparing it with the distance threshold. Among them, the task target point can be the end point or the middle point of the entire task (that is, in the long-distance path-finding process, some points are set between the starting point and the end point of the road so that the end point can be reached smoothly, and these set points are the middle points). When the Euclidean distance between the agent and the task target point is less than the distance threshold, it indicates that the agent is gradually approaching the task target point, and a one-time high reward (as a specific example of a positive reward value) is immediately triggered; otherwise, a negative reward value or no reward is given.
[0051] The specific value of the high reward or negative reward can be pre-set, or determined by the functional relationship between the simulated distance and the reward value, and the embodiment of the present application is not limited thereto. In addition, the distance threshold can be flexibly set according to the actual game task.
[0052] In actual game scenarios, the agent may have to climb slopes, walk up stairs, or take elevators while performing tasks. For this kind of movement with height differences (if the ground in the game is set as the XY plane in the coordinate axis, the height belongs to the Z axis, indicating the distance from the ground), this application proposes a dynamic distance reward, which simultaneously evaluates changes in the XY plane (i.e., horizontal direction) and the Z axis direction (i.e., vertical direction), that is, giving positive or negative rewards (i.e., positive reward values or negative reward values) based on whether the agent is close to the task target point.
[0053] Regarding the dynamic distance reward, in some embodiments of the present application, when there is a height difference between the mission target point and the current position of the agent, S220 may further include:
[0054] S221: If the horizontal distance between the horizontal position of the agent and the task target point is not less than a distance threshold, the reward value is calculated according to a preset rule.
[0055] For example, in some embodiments of the present application, when the horizontal position of the agent's actual position on the XY plane is greater than or equal to a distance threshold, the reward value is calculated according to the three-dimensional Euclidean distance (as a specific example of a preset rule).
[0056] S222: If the horizontal distance between the horizontal position and the task target point is less than a distance threshold and the vertical distance between the agent and the task target point is greater than a height threshold, the reward value is generated in the vertical direction.
[0057] For example, in some embodiments of the present application, when the actual position of the agent after performing the next action after moving in both the XY plane and the Z axis, in the XY plane, the horizontal distance between the ground position of the agent and the horizontal point of the vertical projection of the task target point to the ground is less than a distance threshold (e.g., 20 meters), and the Z-axis height difference (i.e., vertical distance) between the height of the agent on the Z axis and the height of the task target point is greater than a height threshold (e.g., 1 meter), a double reward is given for the change in the Z-axis direction (as a specific example of a reward value). By setting a special Z-axis direction reward, it is ensured that the agent has a high degree of adaptability in complex terrain (such as buildings and platforms).
[0058] S223: If the horizontal distance between the horizontal position and the mission target point is less than a distance threshold and the vertical distance between the vertical height and the mission target point is not greater than a height threshold, then generate the reward value in the horizontal direction.
[0059] For example, in some embodiments of the present application, when the actual position of the agent after moving on both the XY plane and the Z axis after performing the next action, on the XY plane, the horizontal distance between the ground position of the agent and the horizontal point of the vertical projection of the task target point to the ground is less than a distance threshold, and the Z-axis height difference between the height of the agent on the Z axis and the height of the task target point is less than or equal to the height threshold, a reward value is only given for the change on the XY plane.
[0060] In addition, if there is a height difference between the task target point and the actual position of the current agent, and no valid Z-axis reward is obtained within a certain period of time, the time-decay penalty mechanism will be activated (i.e., a negative reward value will be given).
[0061] The reward value generated in the vertical direction is within the set reward range.
[0062] For example, to avoid reward oscillation, set a change filter for the Z-axis reward (such as ignoring it within 0.1 meters) and limit the maximum / minimum rewards (as a specific example of the reward range).
[0063] It should be understood that the distance threshold, height threshold and reward range can be flexibly adjusted according to the actual game scenario, and the embodiments of the present application are not specifically limited here.
[0064] Regarding obstacle rewards, in some embodiments of the present application, if it is confirmed that the distance between the actual position and the target obstacle is less than the distance between the current position of the agent and the target obstacle, a negative reward value is generated; the negative reward value represents that the agent is gradually approaching the target obstacle.
[0065] For example, in some embodiments of the present application, before the intelligent agent performs the next action, it first determines the distance between the current position and the target obstacle. After the intelligent agent performs the next action, it determines the distance between the actual position and the target obstacle. If the distance after the action is performed becomes smaller than before, it indicates that the intelligent agent is approaching the target obstacle and cannot achieve intelligent obstacle avoidance. In this case, a negative reward value is given.
[0066] Alternatively, you can also perform pixel-level detection on the bottom center region of the field of view (ROI) of the instance segmentation map output by the instance segmentation network, and compare the segmented pixel values with the obstacle color calibrated in the pre-configured items to see if they are within the pixel tolerance range (for example, ±10). If so, it indicates that the obstacle in the image information can be accurately segmented, and a positive reward value can be given. Otherwise, it indicates that the obstacle is not accurately segmented, and a negative penalty, that is, a negative reward value, can be given. The ROI area size can be adaptively adjusted according to the resolution (such as 480p or 240p). ROI bottom obstacle recognition is close to the human field of view model and can significantly improve navigation robustness.
[0067] In addition, this application can set hyperparameters to control the variation range of reward values to prevent outliers; at the same time, the variation range of reward values is controlled within an order of magnitude to prevent the strategy from being dominated by a single item.
[0068] S230, adjusting the network parameters of the fusion network based on the reward value to obtain a target fusion model; wherein the target fusion model is used to provide action and direction navigation for the agent when the agent performs a game task.
[0069] For example, in some embodiments of the present application, the reward values obtained in the above embodiments are used to optimize the parameters of the fusion network to obtain a target fusion model that can accurately provide action and direction navigation for the intelligent agent.
[0070] When executing game tasks, by inputting the image features corresponding to the image information displayed on the current game interface and the posture vector features of the intelligent body into the target fusion model, the next action and direction of the intelligent body can be output, thereby achieving accurate navigation of the intelligent body and improving the efficiency of completing game tasks.
[0071] The following is combined with Figure 3 The specific process of mapless navigation based on an open 3D world provided by some embodiments of the present application is exemplified.
[0072] Please see the attached Figure 3 , Figure 3 A flowchart of a method for mapless navigation based on an open 3D world is provided for some embodiments of the present application.
[0073] The above process is described below as an example.
[0074] S310, capturing image information of the game interface currently displayed on the screen.
[0075] S320: Input the image information into the instance segmentation network and the monocular depth network respectively, and output the image features.
[0076] S330, obtaining the posture vector feature corresponding to the posture information of the intelligent body.
[0077] S340, input the image features and pose vector features into the fusion network to predict the next action of the intelligent agent.
[0078] S350, determining a reward value by analyzing the actual position of the agent after performing the next action.
[0079] S360: Adjust the network parameters of the fusion network based on the reward value to obtain a target fusion model.
[0080] It should be noted that S320 and S330 can be executed simultaneously, or S330 can be executed first and then S320, which is not specifically limited in this application. The specific implementation process of S310 to S360 can refer to the method embodiment provided above. To avoid repetition, detailed description is appropriately omitted here.
[0081] As can be seen from the above embodiments of the present application, the present application has the following advantages over the traditional method:
[0082] 1) Get rid of dependence on static maps and improve adaptability to dynamic environments;
[0083] Traditional methods rely on NavMesh or global map planning and cannot cope with real-time obstacle changes, resulting in poor system robustness. The present invention replaces the static navigation map with a dynamic visual input mechanism based on instance segmentation maps and depth estimation maps; and through semantic detection of the bottom ROI area, it accurately detects and punishes potential obstacle collision areas in each step. This design combines a dynamic distance reward function with an obstacle penalty mechanism (i.e., obstacle reward) to enable the system to actively adapt to dynamic environments. Experimental results show that even when encountering interference from high-speed vehicles on urban roads, the intelligent agent can still maintain its navigation direction, which is significantly better than traditional static map navigation strategies.
[0084] 2) The cross-floor three-dimensional navigation capability is significantly enhanced;
[0085] Traditional reinforcement learning strategies lack explicit Z-axis rewards, which can lead to a loss of ability to explore upper and lower floors in multi-story structures. This invention effectively guides paths with height differences through a dedicated Z-axis reward mechanism and a distance-related weighting control strategy. The system accurately identifies paths with height differences, such as stairs and elevators, and automatically triggers a penalty mechanism if the Z-axis reward interval exceeds 30 seconds when there is a height difference, preventing the strategy from remaining on the same plane for extended periods. In actual testing, the agent achieved a completion rate of 75% in indoor building navigation tasks requiring multiple floor crossings, significantly outperforming traditional algorithms that lack height perception mechanisms.
[0086] 3) Improve the stability and generalization ability of navigation strategies
[0087] By introducing a multimodal network structure, the present invention can make full use of the image processing branch and vector branch information: the image processing branch uses a four-level progressive convolutional network to extract multi-scale semantic features, and the vector branch extracts environmental information such as the current state, orientation, and obstacle layout. The fused features are modeled in the time dimension through LSTM (as a specific example of a sub-model within the fusion network), and the output is jointly decided by the policy branch and the value branch. The fusion network optimizes policy updates through the PPO algorithm, effectively avoiding the problem of policy oscillation. Comparative experiments show that the network structure exhibits high robustness and policy stability in multiple different terrain tasks, good transferability, and faster training convergence.
[0088] 4) Fusion of relative depth maps and instance segmentation maps to overcome training limitations;
[0089] The combined input of relative depth maps and instance segmentation maps fully integrates spatial geometry and semantic structure information, enabling more accurate identification of traversable areas. Multi-scale fusion is achieved by training a more stable image feature extraction network, effectively enhancing strategic perception and generalization capabilities across floors, in complex terrain, and in obscured environments. The integration of relative depth maps and instance segmentation maps is a key supporting technology for achieving highly robust 3D navigation. Through dual-modal input, the model not only achieves geometric and semantic complementarity, but also effectively overcomes the cognitive limitations and training bottlenecks brought about by a single perception source, significantly improving overall task completion rate and stability.
[0090] In summary, this application effectively solves the key technical shortcomings of existing technologies in dynamic obstacle avoidance, complex path planning, three-dimensional structure understanding, etc. through systematic technical improvements in reward mechanism design, perception structure optimization and spatiotemporal information fusion, and realizes repeatable, high-success-rate and high-robustness navigation control tasks in complex environment navigation scenarios, with significant technological advancement and application promotion value.
[0091] Please refer to Figure 4 , Figure 4 A block diagram illustrating the components of an apparatus for open 3D world mapless navigation provided by some embodiments of the present application is provided. It should be understood that the apparatus for open 3D world mapless navigation corresponds to the aforementioned method embodiments and is capable of executing each of the steps involved in the aforementioned method embodiments. The specific functions of the apparatus for open 3D world mapless navigation can be found in the description above, and a detailed description is omitted here to avoid repetition.
[0092] Figure 4The device for mapless navigation in an open 3D world includes at least one software functional module that can be stored in a memory or fixed in the device in the form of software or firmware. The device for mapless navigation in an open 3D world includes: a fusion module 410 for inputting image features of a game interface currently displayed on the screen and pose vector features of an agent into a fusion network to predict the agent's next action; wherein the image features include object information and relative distances between objects; and the next action includes the agent's action type and movement direction; a reward module 420 for determining a reward value by analyzing the actual position of the agent after executing the next action; wherein the reward value includes a positive reward value and a negative reward value; the positive reward value indicates that the agent is gradually approaching the task target point, and the negative reward value indicates that the agent is gradually moving away from the task target point or approaching a target obstacle; and an optimization module 430 for adjusting network parameters of the fusion network based on the reward value to obtain a target fusion model; wherein the target fusion model is used to provide action and direction navigation for the agent when the agent performs a game task.
[0093] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.
[0094] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0095] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0096] like Figure 5 As shown, some embodiments of the present application provide an electronic device 500, which includes: a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520, wherein the processor 520 can implement a method as described in any of the above embodiments when reading the program from the memory 510 through the bus 530 and executing the program.
[0097] Processor 520 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, processor 520 can be a microprocessor.
[0098] The memory 510 can be used to store instructions executed by the processor 520 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all functions of one or more modules described in the embodiments of this application. The processor 520 of the embodiment of the present disclosure can be used to execute the instructions in the memory 510 to implement the method shown above. The memory 510 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memory known to those skilled in the art.
[0099] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0100] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0101] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
Claims
1. A method for mapless navigation based on an open 3D world, characterized in that: include: Inputting the image features of the game interface currently displayed on the screen and the pose vector features of the agent into a fusion network to predict the next action of the agent; wherein the image features include: object information and relative distances between objects; the next action includes: the action type and movement direction of the agent; Determining a reward value by analyzing the actual position of the agent after performing the next action; wherein the reward value includes a positive reward value and a negative reward value; the positive reward value indicates that the agent is gradually approaching the task target point, and the negative reward value indicates that the agent is gradually moving away from the task target point or gradually approaching the target obstacle; The network parameters of the fusion network are adjusted based on the reward value to obtain a target fusion model; wherein the target fusion model is used to provide action and direction navigation for the agent when the agent performs a game task.
2. The method according to claim 1, wherein Before inputting the image features of the game interface currently displayed on the screen and the pose vector features of the agent into the fusion network, the method further includes: Obtaining image information of the game interface displayed on the current screen, and obtaining position information of the agent on the current screen; Analyzing the image information to determine image features in the image information; Generate the posture vector feature corresponding to the posture information.
3. The method according to claim 2, wherein The image features include: an instance segmentation map and a relative depth map; and analyzing the image information to determine the image features in the image information includes: Inputting the image information into a pre-trained instance segmentation network to obtain the instance segmentation map; wherein the instance segmentation map represents the object information of various objects contained in the image information; The image information is input into a pre-trained monocular depth network to obtain the relative depth map; wherein the relative depth map represents the relative distances between various objects.
4. The method according to any one of claims 1 to 3, wherein The determining of the reward value by analyzing the actual position of the agent after performing the next action includes: When it is determined that the relative distance between the actual position and the task target point is less than a distance threshold, the positive reward value is triggered to be generated.
5. The method according to any one of claims 1 to 3, wherein When there is a height difference between the mission target point and the current position of the agent, the actual position includes a horizontal position and a vertical height from the ground; The determining of the reward value by analyzing the actual position of the agent after performing the next action includes: If the horizontal distance between the horizontal position of the agent and the task target point is not less than a distance threshold, the reward value is calculated according to a preset rule; If the horizontal distance between the horizontal position and the task target point is less than a distance threshold and the vertical distance between the agent and the task target point is greater than a height threshold, generating the reward value in the vertical direction; If the horizontal distance between the horizontal position and the mission target point is less than a distance threshold and the vertical distance between the vertical height and the mission target point is not greater than a height threshold, the reward value is generated in the horizontal direction.
6. The method according to claim 5, wherein The reward value generated in the vertical direction is within a set reward range.
7. The method according to any one of claims 1 to 3, wherein The determining of the reward value by analyzing the actual position of the agent after performing the next action includes: If it is confirmed that the distance between the actual position and the target obstacle is less than the distance between the current position of the agent and the target obstacle, a negative reward value is generated; the negative reward value represents that the agent is gradually approaching the target obstacle.
8. A device for mapless navigation based on an open 3D world, characterized in that: include: A fusion module is configured to input the image features of the game interface currently displayed on the screen and the pose vector features of the agent into a fusion network to predict the next action of the agent; wherein the image features include: object information and relative distances between objects; the next action includes: the action type and movement direction of the agent; A reward module is configured to determine a reward value by analyzing the actual position of the agent after performing the next action; wherein the reward value includes a positive reward value and a negative reward value; the positive reward value indicates that the agent is gradually approaching the task target point, and the negative reward value indicates that the agent is gradually moving away from the task target point or gradually approaching a target obstacle; An optimization module is used to adjust the network parameters of the fusion network based on the reward value to obtain a target fusion model; wherein the target fusion model is used to provide action and direction navigation for the agent when the agent performs a game task.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 7 when the processor runs the computer program.
Citation Information
Patent Citations
Player analysis using one or more neural networks
CN114375218A
Content recommendations using one or more neural networks
US20210064965A1