Intelligent agent navigation method based on dynamic obstacle detection
By acquiring RGB images and depth images around the agent, identifying obstacle areas, building envelope cubes, calculating exploration coverage, combining obstacle type and relative distance, and dynamically adjusting obstacle avoidance strategies, the problem that the agent navigation method in the existing technology cannot respond to dynamic environmental changes in real time, achieving more efficient and accurate navigation effects.
Patent Information
- Application Number
- CN202510575242.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the agent navigation method relies on the feasible path data pre-generated by the game server and cannot respond to dynamic environment changes in real time, resulting in inflexible and inaccurate navigation.
By acquiring RGB images and depth images around the agent, identifying obstacle areas, building envelope cubes, calculating exploration coverage, combining obstacle type and relative distance, dynamically adjusting obstacle avoidance strategies, and setting a reward and punishment mechanism to improve navigation accuracy and efficiency.
It realizes precise navigation of the agent in a dynamic environment, improves navigation flexibility and security, enhances adaptability to different environments, and improves target proximity and exploration efficiency.
Smart Images

Figure CN120393423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to an agent navigation method based on dynamic obstacle detection. Background Art
[0002] In open-world 3D games, the automatic navigation of agents (such as NPCs or player characters) is a key function that affects the playability and user experience of the game. However, existing navigation technologies have some limitations and are difficult to meet the requirements of increasingly complex and dynamic game scenarios.
[0003] The patent document with the publication number CN106060052A discloses a three-dimensional navigation method for mobile terminal online games. After the mobile terminal logs in to the game, it determines whether the user's game character uses the three-dimensional navigation function to walk automatically. If so, it sends a navigation request to the game server. The game server generates feasible path data according to the navigation request and transmits the feasible path data to the game program module. The game program module controls the game character to walk automatically along the feasible path according to the feasible path data.
[0004] It can be seen that the following problems exist: In the prior art, due to relying on the feasible path data pre-generated by the game server, this static path planning cannot respond to dynamic changes in the environment in real time. Summary of the Invention
[0005] For this reason, the present invention provides an agent navigation method based on dynamic obstacle detection to overcome the problem that in the prior art, relying on the feasible path data pre-generated by the game server, this static path planning cannot respond to dynamic changes in the environment in real time.
[0006] To achieve the above object, the present invention provides an agent navigation method based on dynamic obstacle detection, including:
[0007] Step S1, obtaining image data of the environment around the agent, where the image data includes RGB images and depth images;
[0008] Step S2, identifying the obstacle area in the image according to the RGB image and the depth image, determining the initial position of the agent, and determining the relative distance between the agent and the obstacle according to the obstacle and the initial position;
[0009] Step S3, constructing an envelope cube according to the initial position to obtain the exploration path of the agent, and calculating the coverage rate according to the envelope cube and the exploration path to obtain the exploration coverage rate;
[0010] Step S4, perform feature extraction on the obstacle area to determine the obstacle type of the obstacle, and calculate the influence degree of the obstacle on the agent according to the obstacle type, the exploration coverage rate, and the relative distance;
[0011] Step S5, determine the obstacle avoidance strategy according to the influence degree, and adjust the relative distance according to the obstacle avoidance strategy to obtain the distance between the agent and the preset target position, so as to obtain the target proximity;
[0012] Step S6, determine the reward and punishment mechanism according to the target proximity and the influence degree, calculate the exploration time of the agent according to the exploration path and the exploration coverage rate, compare the exploration time with a preset exploration time threshold to obtain a comparison result, and adjust the reward and punishment mechanism according to the comparison result and the target proximity.
[0013] Further, the process of identifying the obstacle area in the image according to the RGB image and the depth image and determining the initial position of the agent includes:
[0014] Construct a four-layer progressive convolutional network according to the RGB image to obtain multi-dimensional image features;
[0015] Input the multi-dimensional image features into a pre-trained instance segmentation network to identify and segment each obstacle area of the RGB image;
[0016] Input the depth image into a pre-trained monocular depth estimation network to obtain a corresponding relative depth map;
[0017] Use a feature point matching algorithm to extract the matching feature points of the RGB image and the depth image to obtain matching feature points;
[0018] Determine the initial position of the agent according to the relative depth map and the matching feature points.
[0019] Further, the process of determining the relative distance between the agent and the obstacle according to the obstacle and the initial position includes:
[0020] Determine the three-dimensional position and size of each obstacle according to the obstacle and the relative depth map to obtain the obstacle position and the obstacle size;
[0021] Calculate the relative distance according to the initial position, the obstacle position, and the obstacle size.
[0022] Further, the process of step S3 includes:
[0023] Construct a three-dimensional coordinate system centered on the initial position, and construct an enveloping cube according to the three-dimensional coordinate system. The size of the enveloping cube covers the expected exploration area of the agent;
[0024] Obtain the maximum speed, acceleration, and turning radius of the agent to evaluate the motion ability of the agent;
[0025] Determine the exploration path according to the motion ability, the enveloping cube, and the distribution information of the obstacles;
[0026] Divide the enveloping cube into a number of equal-sized volume units. When the agent moves along the exploration path, dynamically mark the volume units passed by the agent;
[0027] Calculate the exploration coverage rate according to the ratio of the number of volume units passed by the agent to the total number of volume units of the enveloping cube.
[0028] Further, the process of calculating the exploration coverage rate according to the ratio of the number of volume units passed by the agent to the total number of volume units of the enveloping cube includes:
[0029] Create a three-dimensional array, or three-dimensional matrix, to record whether each volume unit has been explored by the agent. If the volume unit has been explored, obtain the trajectory point of the agent;
[0030] When the trajectory point is mapped to a certain volume unit, mark the exploration state of the volume unit as explored;
[0031] Traverse the three-dimensional array, or the three-dimensional matrix, and count the number of volume units with the exploration state of explored;
[0032] Calculate the total number of volume units according to the size of the enveloping cube;
[0033] Obtain the exploration coverage rate according to the ratio of the number of volume units and the total number of volume units.
[0034] Further, the process of extracting features from the obstacle area to determine the obstacle type of the obstacle includes:
[0035] Extract features of the shape, color, texture, and depth of the obstacle to obtain multiple feature results;
[0036] Fuse multiple feature results to obtain a multi-dimensional feature vector;
[0037] Input the multi-dimensional feature vector into a trained classifier to obtain the obstacle type.
[0038] Further, the process of calculating the influence degree of the obstacle on the agent according to the obstacle type, the exploration coverage rate, and the relative distance includes:
[0039] Assign a weight to each obstacle type to obtain an assignment result;
[0040] Analyze the influence degree of the exploration coverage rate on the obstacle to obtain a first analysis result;
[0041] Analyze the influence degree of the relative distance on the obstacle to obtain a second analysis result;
[0042] Calculate the comprehensive influence degree according to the assignment result, the first analysis result, and the second analysis result to obtain the influence degree.
[0043] Further, the process of step S5 includes:
[0044] Set thresholds corresponding to different influence degrees to obtain a first degree threshold, a second degree threshold, and a third degree threshold;
[0045] Compare according to the first degree threshold, the second degree threshold, the third degree threshold, and the influence degree to obtain a corresponding obstacle avoidance strategy;
[0046] Adjust the relative distance according to the obstacle avoidance strategy to obtain a strategy adjustment distance;
[0047] Use the Euclidean distance to calculate the relative distance between the strategy adjustment distance and the target distance to obtain the target proximity.
[0048] Further, the process of determining the reward and punishment mechanism according to the target proximity and the influence degree includes:
[0049] Set a first target proximity threshold, a second target threshold, and a third target threshold, and compare according to the first target proximity threshold, the second target threshold, the third target threshold, and the target proximity to obtain a comparison result;
[0050] Assign a score to the first degree threshold, the second degree threshold, and the third degree threshold to obtain a scoring result;
[0051] Determine the reward and punishment mechanism according to the comparison result and the scoring result.
[0052] Further, the process of comparing the exploration time with a preset exploration time threshold to obtain a comparison result, and adjusting the reward and punishment mechanism according to the comparison result and the target proximity includes:
[0053] If the actual exploration time is less than or equal to the exploration time threshold, the exploration time is considered reasonable;
[0054] If the actual exploration time is greater than the exploration time threshold, the exploration time is considered unreasonable;
[0055] Among them, when the exploration time is reasonable, when the target proximity is less than or equal to the first target threshold, a high-level reward is given. When the target proximity is greater than the first target proximity threshold and less than or equal to the second target threshold, a medium-level reward is given. When the target proximity is greater than the second target proximity threshold, a low-level reward is given;
[0056] When the exploration time is unreasonable, when the target proximity is greater than or equal to the second target threshold, a mild penalty is given. When the target proximity is less than the second target threshold, a high penalty is given.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows. By acquiring the RGB image and depth image of the environment around the agent and identifying the obstacle area in the image, the relative distance between the agent and the obstacle can be accurately determined, thereby improving the accuracy of navigation. Constructing an envelope cube and calculating the exploration coverage rate enables the agent to dynamically adjust its exploration path according to the changes in the environment, enhancing the agent's adaptability to different environments. By extracting the features of the obstacle area and determining the obstacle type, combining the exploration coverage rate and the relative distance to calculate the influence degree of the obstacle on the agent, a better obstacle avoidance strategy can be formulated to improve the safety of the agent in a complex environment. Determining the obstacle avoidance strategy according to the influence degree and adjusting the relative distance enables the agent to approach the target position more efficiently, improving the target proximity. Determining the reward and punishment mechanism according to the target proximity and the influence degree, and comparing and adjusting the reward and punishment mechanism according to the exploration time and the preset exploration time threshold, encourages the agent to adopt a better navigation strategy and improves the exploration efficiency. The clear reward and punishment mechanism provides effective feedback for the learning algorithm of the agent, accelerates the learning process, and improves the autonomous learning ability of the agent.
[0058] In particular, by constructing a three-dimensional coordinate system centered on the initial position and constructing an envelope cube, it is ensured that the expected exploration area of the agent is completely covered. Obtaining the motion ability parameters of the agent such as the maximum speed, acceleration, and turning radius provides an important basis for determining the exploration path. According to these parameters and the obstacle distribution information, the exploration path can be optimized to improve the exploration efficiency and reduce the collision risk. When the agent moves along the exploration path, it dynamically marks the volume units passed through, realizing real-time feedback and monitoring.
[0059] In particular, by calculating the exploration coverage rate, unexplored areas can be identified, enabling targeted resource allocation, improving exploration efficiency, and avoiding repeated exploration of known areas. The calculation results of the exploration coverage rate can provide feedback to the path planning algorithm, helping the agent select a more effective exploration path and reducing unnecessary movement and energy consumption.
[0060] In particular, by assigning different weights to different types of obstacles, the agent can more accurately evaluate the impact of each obstacle on its travel path, thus making more reasonable obstacle avoidance decisions. Through a multi-dimensional analysis of comprehensive weight, exploration coverage rate, and relative distance, the agent can quickly calculate the degree of influence of the obstacle and respond promptly to avoid collisions and delays.
[0061] In particular, by rewarding agents that efficiently complete exploration tasks, their exploration efficiency can be incentivized. Penalizing agents that exceed the exploration time threshold encourages them to avoid inefficient exploration. Dynamically adjusting rewards and penalties based on the actual exploration time and proximity to the goal makes the mechanism more flexible and effective. A clear reward and penalty mechanism provides effective feedback to the agent's learning algorithm, promoting its autonomous learning and optimization process. Brief Description of the Drawings
[0062] Figure 1 It is a schematic flowchart of the agent navigation method based on dynamic obstacle detection provided by an embodiment of the present invention;
[0063] Figure 2 It is a schematic flowchart of the agent navigation method based on dynamic obstacle detection provided by an embodiment of the present invention;
[0064] Figure 3 It is a schematic flowchart of the agent navigation method based on dynamic obstacle detection provided by an embodiment of the present invention;
[0065] Figure 4 It is a schematic flowchart of the agent navigation method based on dynamic obstacle detection provided by an embodiment of the present invention. Detailed Embodiments
[0066] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0067] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0068] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0069] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0070] Please refer to Figure 1 As shown, an intelligent agent navigation method based on dynamic obstacle detection provided in this embodiment includes:
[0071] Step S1, obtaining image data of the environment around the intelligent agent, where the image data includes RGB images and depth images;
[0072] Step S2, identifying the obstacle area in the image according to the RGB image and the depth image, determining the initial position of the intelligent agent, and determining the relative distance between the intelligent agent and the obstacle according to the obstacle and the initial position;
[0073] Step S3, constructing an envelope cube according to the initial position to obtain the exploration path of the intelligent agent, and calculating the coverage rate according to the envelope cube and the exploration path to obtain the exploration coverage rate;
[0074] Step S4, extracting features from the obstacle area to determine the obstacle type, and calculating the influence degree of the obstacle on the intelligent agent according to the obstacle type, the exploration coverage rate, and the relative distance;
[0075] Step S5, determining an obstacle avoidance strategy according to the influence degree, and adjusting the relative distance according to the obstacle avoidance strategy to obtain the distance between the intelligent agent and a preset target position, so as to obtain the target proximity;
[0076] Step S6: Determine the reward and punishment mechanism based on the target proximity and the degree of influence. Calculate the exploration time of the agent according to the exploration path and the exploration coverage rate. Compare the exploration time with a preset exploration time threshold to obtain a comparison result, and adjust the reward and punishment mechanism according to the comparison result and the target proximity.
[0077] Specifically, the camera carried by the agent captures both RGB images and depth images simultaneously. The RGB images are used to identify the color and texture features of obstacles. The depth images are used to measure the distance between the obstacles and the agent. Potential obstacle regions are identified from the RGB images using image processing algorithms (such as edge detection and region growing). Combining the depth images, the actual positions and sizes of these regions are determined through depth measurement. Calculate the relative distance between the initial position of the agent and each obstacle. With the initial position of the agent as the center, construct a virtual envelope cube that defines the exploration space of the agent. Calculate the proportion of the space covered by the agent during exploration, i.e., the exploration coverage rate, according to the envelope cube and the preset exploration path. Feature extraction is performed on each identified obstacle region, including shape, color, texture, and depth features. The extracted features are fused into a multi-dimensional feature vector and input into a pre-trained classifier to determine the obstacle type (dynamic, static, spatial obstacle). Combining the obstacle type, the exploration coverage rate, and the relative distance, calculate the degree of influence of each obstacle on the agent. For example, dynamic obstacles may have a higher influence weight. Select an appropriate obstacle avoidance strategy according to the degree of influence of the obstacle, such as detouring, stopping, or decelerating. Adjust the travel path of the agent to avoid obstacles, and at the same time calculate the new relative distance. According to the adjusted path, calculate the distance between the agent and the preset target position, i.e., the target proximity. Set an initial reward and punishment mechanism based on the target proximity and the degree of influence of the obstacles. For example, a reward is obtained for successfully avoiding an obstacle, and a punishment is received for a collision. Calculate the time required for the agent to complete the exploration path and compare it with the preset exploration time threshold. Dynamically adjust the reward and punishment mechanism according to the comparison result and the target proximity. For example, if the exploration time is too long or the target proximity is insufficient, increase the punishment intensity.
[0078] Specifically, by obtaining the RGB image and depth image of the environment around the agent and identifying the obstacle regions in the images, the relative distance between the agent and the obstacles can be accurately determined, thereby improving the accuracy of navigation. Constructing an envelope cube and calculating the exploration coverage rate enables the agent to dynamically adjust its exploration path according to changes in the environment, enhancing the agent's adaptability to different environments. By extracting features from the obstacle regions and determining the obstacle types, and combining the exploration coverage rate and relative distance to calculate the impact degree of the obstacles on the agent, a better obstacle avoidance strategy can be formulated to improve the safety of the agent in complex environments. Determining the obstacle avoidance strategy according to the impact degree and adjusting the relative distance enables the agent to approach the target position more efficiently, improving the target proximity. Determining the reward and punishment mechanism according to the target proximity and impact degree, and comparing and adjusting the reward and punishment mechanism according to the exploration time and the preset exploration time threshold, encourages the agent to adopt a better navigation strategy and improves the exploration efficiency. The clear reward and punishment mechanism provides effective feedback for the learning algorithm of the agent, accelerates the learning process, and improves the autonomous learning ability of the agent.
[0079] Specifically, the process of identifying the obstacle regions in the images according to the RGB image and the depth image and determining the initial position of the agent includes:
[0080] Constructing a four-layer progressive convolutional network according to the RGB image to obtain multi-dimensional image features;
[0081] Inputting the multi-dimensional image features into a pre-trained instance segmentation network to identify and segment each obstacle region of the RGB image;
[0082] Inputting the depth image into a pre-trained monocular depth estimation network to obtain the corresponding relative depth map;
[0083] Using a feature point matching algorithm to extract the matching feature points of the RGB image and the depth image to obtain matching feature points;
[0084] Determining the initial position of the agent according to the relative depth map and the matching feature points.
[0085] Specifically, a four-level progressive convolutional layer structure (Conv2d with 8→16→32→64 channels) is adopted, and each level of convolutional layer is followed by a LayerNorm normalization layer and a ReLU activation function. The convolution stride is designed to be 2, gradually compressing the size of the feature map (120×160→60×80→30×40→15×20→8×10). Finally, it is unfolded into a 5120-dimensional feature vector through a Flatten layer to obtain a multi-dimensional image feature representation. The integrated image features are input into a pre-trained instance segmentation network. The network analyzes the features, identifies and segments each obstacle region in the image, such as pedestrians, vehicles, etc. The captured depth image is input into the processing system. A pre-trained monocular depth estimation network is used to analyze the depth image and generate a corresponding relative depth map representing the depth information of each point. Significant feature points, such as corner points and edge points, are extracted from the RGB image and the depth image respectively. A feature point matching algorithm, such as RANSAC or FLANN, is used to find the corresponding matching feature point pairs in the RGB image and the depth image. According to the matching feature points and the relative depth map, the three-dimensional spatial coordinates of each matching feature point are calculated. Using the calculated three-dimensional coordinates and geometric relationships, methods such as triangulation or multilateration are used to estimate the initial position of the agent in the environment.
[0086] Specifically, by constructing a four-layer progressive convolutional network, image features from simple to complex can be extracted layer by layer, ensuring that the obstacle features in the RGB image are fully captured. Using a pre-trained instance segmentation network, each obstacle region in the RGB image can be accurately identified and segmented, reducing the cases of misidentification and missed identification. Accurate obstacle identification and initial position determination enable the agent to complete the navigation task more quickly and safely. It reduces the problems of navigation failure or low efficiency caused by incorrect obstacle identification or inaccurate initial position.
[0087] Specifically, the process of determining the relative distance between the agent and the obstacle according to the obstacle and the initial position includes:
[0088] Determine the three-dimensional position and size of each obstacle according to the obstacle and the relative depth map to obtain the obstacle position and the obstacle size;
[0089] Calculate the relative distance according to the initial position, the obstacle position and the obstacle size.
[0090] Specifically, during the game development stage, mark the obstacles in the scene and assign a unique identifier to each obstacle. For dynamically generated obstacles (such as enemies, moving vehicles, etc.), perform real-time marking during generation. Use the rendering system of the game engine or a dedicated depth sensor simulation component to generate a relative depth map containing the depth information of the obstacles. Each pixel value in the depth map represents the distance from the camera (or the perspective of the agent) to the corresponding scene point. For each marked obstacle, find the corresponding area in the relative depth map through its identifier. Utilize the information in the depth map and combine the camera internal parameters (such as focal length, field of view, etc.) to calculate the three-dimensional position (x, y, z) of the obstacle through triangulation. Estimate the size (Width, Height, Depth) of the obstacle based on the pixel coverage area of the obstacle in the depth map. For obstacles with irregular shapes, an approximate cuboid or cylinder can be used to represent them. Obtain the initial position (X_player, Y_player, Z_player) of the agent. For each obstacle, calculate the Euclidean distance (Distance) between its center point and the agent's position:
[0091] Distance = sqrt((x - X player ) 2 +(y - Y player ) 2 +(z - Z player ) 2 )
[0092] Considering the size of the obstacle, calculate the minimum distance (Mindistance):
[0093]
[0094] If the obstacle has an irregular shape, a collision detection algorithm (such as the GJK algorithm) can be further used to accurately calculate the minimum distance.
[0095] Specifically, by accurately calculating the three-dimensional position and size of each obstacle, the agent can more accurately judge the relative distance from the obstacle, thereby improving the accuracy of obstacle avoidance and reducing the collision risk. This method allows the agent to quickly adapt to environments with different complexities and ensure good obstacle avoidance performance in a dynamically changing environment by updating the obstacle information in real time.
[0096] Specifically, as Figure 2 shown, the process of step S3 includes:
[0097] Step S31, construct a three-dimensional coordinate system centered on the initial position, and construct an envelope cube according to the three-dimensional coordinate system. The size of the envelope cube covers the expected exploration area of the agent;
[0098] Step S32, obtaining the maximum speed, acceleration, and turning radius of the agent to evaluate the agent's movement ability;
[0099] Step S33, determining the exploration path according to the movement ability, the envelope cube, and the distribution information of the obstacles;
[0100] Step S34, dividing the envelope cube into a number of volume units of equal size, and dynamically marking the volume units passed by the agent as it moves along the exploration path;
[0101] Step S35 , calculating the exploration coverage rate according to the ratio of the number of volume units passed by the agent to the total number of volume units of the enveloping cube.
[0102] Specifically, establish a three-dimensional rectangular coordinate system with the initial position (x, y, z) as the origin. Set the directions of the coordinate axes, typically aligning them with the coordinate axes of the game world. Calculate the required side lengths of the enveloping cube based on the agent's intended exploration area. Consider factors such as the agent's range of movement, task requirements, and environment size. In the three-dimensional coordinate system, construct a cube with a predetermined side length centered at the initial position. Ensure that the cube completely covers the agent's intended exploration area, leaving no blind spots. Obtain the agent's maximum speed from the agent's properties or the game engine. Consider speed variations under different terrains and conditions. Accurately obtain the acceleration required for the agent to reach its maximum speed from rest. Consider the nonlinear characteristics of acceleration and deceleration. Measure or calculate the minimum turning radius the agent can achieve during movement. Consider the impact of different speeds and terrains on the turning radius. Utilize obstacle information provided by sensors or the game engine to conduct a detailed analysis of obstacle type, location, size, and distribution. Construct an obstacle map for path planning.
[0103] Specifically, by constructing a three-dimensional coordinate system centered on the initial position and building an enveloping cube, the agent's intended exploration area is fully covered. Obtaining the agent's motion parameters, such as maximum speed, acceleration, and turning radius, provides an important basis for determining the exploration path. Based on these parameters and obstacle distribution information, the exploration path can be optimized to improve exploration efficiency and reduce collision risk. As the agent moves along the exploration path, it dynamically marks the volume units it passes through, enabling real-time feedback and monitoring.
[0104] Specifically, the process of calculating the exploration coverage rate based on the ratio of the number of volume units passed by the agent to the total number of volume units of the enveloping cube includes:
[0105] Create a three-dimensional array, or a three-dimensional matrix, to record whether each of the volume units has been explored by the agent. If a volume unit has been explored, obtain the trajectory point of the agent;
[0106] When the trajectory point is mapped to a certain volume unit, mark the exploration status of this volume unit as explored;
[0107] Traverse the three-dimensional array, or the three-dimensional matrix, and count the number of volume units with the exploration status of explored;
[0108] Calculate the total number of volume units according to the size of the bounding cube;
[0109] Obtain the exploration coverage rate according to the ratio of the number of volume units and the total number of volume units.
[0110] Specifically, according to the size of the bounding cube and the size of the volume unit, determine the size of the three-dimensional array. For example, if the side length of the bounding cube is L and the side length of the volume unit is l, then the size of the array is N×N×N, where N = L / l. Create a three-dimensional array of N×N×N, and set all elements to "unexplored" (for example, represented by 0) in the initial state. Obtain the trajectory points from the movement records or sensor data of the agent. For each trajectory point, calculate the index of the volume unit it is in. This can be achieved by dividing the trajectory point coordinates by the side length of the volume unit and taking the integer. Set the array element of the corresponding volume unit to "explored" (for example, represented by 1). Traverse the entire three-dimensional array and count the number of elements with the value of "explored". Record the counted number of explored volume units. Calculate the total volume of the bounding cube according to its size, that is, L×L×L. Divide the total volume by the volume of a single volume unit (l×l×l) to obtain the total number of volume units. Divide the number of explored volume units by the total number of volume units to obtain the exploration coverage rate. Express the exploration coverage rate in percentage or decimal form for easy understanding and comparison.
[0111] Specifically, by calculating the exploration coverage rate, the unexplored areas can be identified, so as to allocate resources targeted, improve the exploration efficiency, and avoid repeated exploration of known areas. The calculation result of the exploration coverage rate can provide feedback for the path planning algorithm, helping the agent select a more effective exploration path and reducing unnecessary movement and energy consumption.
[0112] Specifically, as Figure 3 shown, the process of extracting features from the obstacle area to determine the obstacle type of the obstacle includes:
[0113] Step S41, extract features of shape, color, texture, and depth of the obstacle to obtain multiple feature results;
[0114] Step S42: Fuse multiple feature results to obtain a multi-dimensional feature vector;
[0115] Step S43: Input the multi-dimensional feature vector into a trained classifier to obtain the obstacle type.
[0116] Specifically, use an edge detection algorithm (such as the Canny operator) to extract the edge information of the obstacle area. Calculate the geometric features of the area, such as area, perimeter, circularity, rectangularity, etc. Use the Hough transform to detect shape features such as lines and circles. Convert the image of the obstacle area into the HSV (hue, saturation, value) color space. Calculate the color histogram and count the distribution of each color channel. Extract color moments, such as mean, variance, skewness, etc. Use Gabor filters to extract the texture features of the obstacle area. Calculate the Local Binary Pattern (LBP) features to describe the local texture structure of the image. Extract the Gray-Level Co-Occurrence Matrix (GLCM) features, such as contrast, energy, homogeneity, etc. Extract the depth information of the obstacle area from the depth map. Calculate the depth histogram and count the distribution of depth values. Extract depth features, such as average depth, depth variance, etc. Integrate the extracted shape, color, texture, and depth features into a feature vector. Ensure that the size of the feature vector is consistent for subsequent processing. Use a feature selection algorithm (such as mutual information method, ReliefF algorithm) to remove redundant and irrelevant features. Apply dimensionality reduction techniques (such as Principal Component Analysis PCA, Linear Discriminant Analysis LDA) to reduce the feature dimension and improve the computational efficiency. According to the characteristics of the obstacle type recognition task, select a suitable classifier, such as Support Vector Machine (SVM), Random Forest (RF), Convolutional Neural Network (CNN), etc. Collect an annotated data set containing various obstacle types. Divide the data set into a training set, a validation set, and a test set. Use the training set data to train the classifier model. Adjust the model parameters, use the validation set for cross-validation, and select the optimal model. Input the extracted and fused multi-dimensional feature vector into the trained classifier model. The classifier outputs the type label of the obstacle. If necessary, the probability or confidence corresponding to the obstacle type can be output to obtain the obstacle type.
[0117] Specifically, by comprehensively considering multi-dimensional features such as shape, color, texture, and depth, obstacles can be more comprehensively described, thereby improving the accuracy of recognition. Accurate obstacle type recognition provides more reliable environmental perception information for the agent, which helps to make more reasonable decisions. Improve the obstacle avoidance strategy and enhance the safety and efficiency of the agent.
[0118] Specifically, the process of calculating the influence degree of the obstacle on the agent according to the obstacle type, the exploration coverage rate, and the relative distance includes:
[0119] Assign a weight to each obstacle type to obtain the assignment result;
[0120] Analyze the degree of influence of the exploration coverage rate on the obstacle to obtain the first analysis result;
[0121] Analyze the degree of influence of the relative distance on the obstacle to obtain the second analysis result;
[0122] Calculate the comprehensive influence degree according to the assignment result, the first analysis result and the second analysis result to obtain the influence degree.
[0123] Specifically, obstacle types include dynamic obstacles (such as vehicles, pedestrians, animals, etc.), static obstacles (such as buildings, trees, furniture, etc.) and spatial obstacles (such as stairs, slopes, steps, etc.). Assign their weights: the weight W1 of static obstacles is 0.2, the weight W2 of dynamic obstacles is 0.5, and the weight W3 of spatial obstacles is 0.3. The exploration coverage rate represents the degree of exploration of the surrounding environment by the agent, usually a value between 0 and 1, where 1 means complete exploration and 0 means no exploration. For each obstacle, analyze its degree of influence on the agent according to its type and exploration coverage rate C. Static obstacles: The higher the exploration coverage rate, the lower the degree of influence, because the agent already knows the position and shape of the static obstacle. Dynamic obstacles: The higher the exploration coverage rate, the higher the degree of influence may be, because the position and state of dynamic obstacles may change at any time. Spatial obstacles: The higher the exploration coverage rate, the lower the degree of influence, because the agent already knows the structure and path of the spatial obstacle. The relative distance D represents the distance between the agent and the obstacle. Analyze the influence of the relative distance on the obstacle: The closer the relative distance, the higher the degree of influence of the obstacle on the agent.
[0124] Example of the calculation formula: The influence degree is inversely proportional to the distance, that is
[0125]
[0126] Combining the weight, the influence of the exploration coverage rate and the influence of the relative distance, calculate the comprehensive influence degree:
[0127] For static obstacles:
[0128] For dynamic obstacles:
[0129] For spatial obstacles:
[0130] Based on the above calculation results, the agent comprehensively evaluates the degree of influence of each obstacle on it, and thus makes corresponding navigation and obstacle avoidance decisions.
[0131] Specifically, by assigning different weights to different types of obstacles, the agent can more accurately evaluate the impact of each obstacle on its travel path, and thus make more reasonable obstacle avoidance decisions. Through multi-dimensional analysis of the comprehensive weight, exploration coverage rate, and relative distance, the agent can quickly calculate the degree of influence of the obstacle, quickly respond, and avoid collisions and delays.
[0132] Specifically, as Figure 4 shown, the process of step S5 includes:
[0133] Step S51, setting thresholds corresponding to different degrees of influence to obtain a first degree threshold, a second degree threshold, and a third degree threshold;
[0134] Step S52, comparing according to the first degree threshold, the second degree threshold, the third degree threshold, and the degree of influence to obtain a corresponding obstacle avoidance strategy;
[0135] Step S53, adjusting the relative distance according to the obstacle avoidance strategy to obtain a strategy adjustment distance;
[0136] Step S54, using the Euclidean distance to calculate the relative distance between the strategy adjustment distance and the target distance to obtain the target proximity.
[0137] Specifically, the first degree threshold (T1): represents a mild influence, and the obstacle has a small impact on the agent's travel. The second degree threshold (T2): represents a moderate influence, and the obstacle has a certain impact on the agent's travel and obstacle avoidance measures need to be taken. The third degree threshold (T3): represents a severe influence, and the obstacle has a great impact on the agent's travel and immediate emergency obstacle avoidance measures need to be taken. Specific values are set for each threshold according to historical data, expert experience, and simulation experiments. For example: T1 = 0.3, T2 = 0.6, T3 = 0.9.
[0138] Specifically, if the influence degree ≤ T1, adopt Strategy A (such as maintaining the current speed and direction). If T1 < influence degree ≤ T2, adopt Strategy B (such as decelerating and adjusting the direction to bypass). If T2 < influence degree ≤ T3, adopt Strategy C (such as emergency braking and significantly adjusting the direction). If the influence degree > T3, adopt Strategy D (such as stopping immediately and backing up). Adjust the distance according to the obstacle avoidance strategy: Strategy A: Do not adjust the relative distance. Strategy B: Increase the relative distance to provide enough space for bypassing. Strategy C: Significantly increase the relative distance to ensure safe obstacle avoidance. Strategy D: Back up to a safe distance. Set the parameters for adjusting the distance, such as the increased distance value or the backed-up distance value. According to the selected strategy, calculate the adjusted relative distance to obtain the strategy adjustment distance. Obtain the target distance: Obtain the distance between the agent and the target from the environmental information or path planning. Use the Euclidean distance calculation: Calculate the relative distance between the strategy adjustment distance and the target distance. The calculation formula for the target proximity P is
[0139] Target proximity = |Strategy adjustment distance - Target distance|
[0140] Specifically, by setting the thresholds for different influence degrees, hierarchical decision-making is achieved, making the obstacle avoidance strategy more flexible and accurate. Dynamically adjust the relative distance according to the influence degree to ensure that the agent maintains the optimal path during obstacle avoidance. Use the Euclidean distance to calculate the target proximity, providing accurate distance information, which helps the agent better plan the travel route. By comprehensively considering the influence degree and the threshold, multiple obstacle avoidance strategies are formulated, effectively improving the safety of the agent in a complex environment.
[0141] Specifically, the process of determining the reward and punishment mechanism according to the target proximity and the influence degree includes:
[0142] Set the first target proximity threshold, the second target threshold, and the third target threshold, and compare them with the target proximity according to the first target proximity threshold, the second target threshold, the third target threshold to obtain a comparison result;
[0143] Assign a score to the first degree threshold, the second degree threshold, and the third degree threshold to obtain a scoring result;
[0144] Determine the reward and punishment mechanism according to the comparison result and the scoring result.
[0145] Specifically, the first target proximity threshold P1 indicates that the target is far away, and the agent has enough space for obstacle avoidance and path adjustment. The second target proximity threshold P2 indicates that the target is at a moderate distance, and the agent needs to start paying attention to obstacle avoidance and path planning. The third target proximity threshold P3 indicates that the target is very close, and the agent needs to precisely avoid obstacles and quickly reach the target. Set P1 to 0.3, P2 to 0.6, and P3 to 0.8 to divide different proximity intervals. For example, if P is greater than P1, the result is "long distance". If P1 is less than or equal to P and greater than P2, the result is "medium distance". If P2 is less than or equal to P and greater than P3, the result is "short distance". If P is less than or equal to P3, the result is "very short distance". Assign a score or penalty value to each impact level interval, such as a low score (or high penalty) for a high impact level and a high score for a low impact level. Combine the target proximity score and the impact level score to determine the final reward or penalty value. A simple addition or multiplication model can be used, or a more complex fusion strategy such as weighted average or neural network can be adopted. Reward and penalty rules: Set specific reward and penalty rules. For example: If the target proximity is high and the impact level is low, give a large reward. If the target proximity is low and the impact level is high, give a large penalty. In other cases, give appropriate rewards or penalties according to the specific score. Dynamically adjust the reward and penalty mechanism according to the real-time performance of the agent and environmental changes. For example, if the agent still cannot effectively avoid high-impact obstacles after multiple attempts, the penalty intensity can be increased to enhance the learning effect.
[0146] Specifically, through the reward and penalty mechanism, the agent is motivated to adopt better obstacle avoidance and path planning strategies. The flexible setting of the threshold and score enables this method to adapt to different environments and scenarios. The clear reward and penalty mechanism provides effective feedback for the learning algorithm of the agent, accelerating the learning process. By punishing high-risk behaviors, the safety of the agent in complex environments is enhanced.
[0147] Specifically, the process of comparing the exploration time with a preset exploration time threshold to obtain a comparison result and adjusting the reward and penalty mechanism according to the comparison result and the target proximity includes:
[0148] If the actual exploration time is less than or equal to the exploration time threshold, the exploration time is considered reasonable;
[0149] If the actual exploration time is greater than the exploration time threshold, the exploration time is considered unreasonable;
[0150] Among them, when the exploration time is reasonable, when the target proximity is less than or equal to the first target threshold, a high reward is given; when the target proximity is greater than the first target proximity threshold and less than or equal to the second target threshold, a medium reward is given; when the target proximity is greater than the second target proximity threshold, a low reward is given;
[0151] When the exploration time is unreasonable, if the target proximity is greater than or equal to the second target threshold, a mild penalty is given; when the target proximity is less than the second target threshold, a severe penalty is given.
[0152] Specifically, a preset exploration time threshold (TT) is set, which represents the maximum time allowed for the agent to complete the exploration task in the game. Specific values are set for the exploration time threshold according to the game difficulty, map size, and agent performance. For example: TT = 300 seconds. When the agent starts the exploration task, a timer is started to record the actual exploration time (AT). When the agent completes the exploration task or reaches a certain specific condition, the timing is stopped and the actual exploration time is recorded. If AT ≤ TT, the exploration time is considered reasonable. If AT > TT, the exploration time is considered unreasonable. The reward and punishment mechanism is adjusted according to the target proximity and the rationality of the exploration time. When the exploration time is reasonable: If the target proximity ≤ the first target proximity threshold (P1): A high reward is given. For example: +50 points. If the first target proximity threshold (P1) < target proximity ≤ the second target proximity threshold (P2): A medium reward is given. For example: +20 points. If the target proximity > the second target proximity threshold (P2): A low reward is given. For example: +10 points. When the exploration time is unreasonable: If the target proximity is greater than or equal to the second target proximity threshold (P2): A mild penalty is given. For example: -10 points. If the target proximity is less than the second target proximity threshold (P2): A severe penalty is given. For example: -30 points. According to the above judgments and the set reward and punishment rules, the score of the agent is adjusted accordingly. The reward and punishment results are fed back to the learning algorithm of the agent for optimizing its subsequent behavior.
[0153] Specifically, by rewarding agents that efficiently complete the exploration task, it motivates them to improve their exploration efficiency. Penalties are imposed on agents that exceed the exploration time threshold to encourage them to avoid inefficient exploration. The reward and punishment are dynamically adjusted according to the actual exploration time and the target proximity, making the mechanism more flexible and effective. The clear reward and punishment mechanism provides effective feedback for the learning algorithm of the agent, promoting its autonomous learning and optimization process.
[0154] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0155] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention; for those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent agent navigation method based on dynamic obstacle detection, characterized in that, Including: Step S1: Obtain image data of the environment around the agent, where the image data includes RGB images and depth images; Step S2: Identify the obstacle regions in the images based on the RGB images and the depth images, and determine the initial position of the agent. Determine the relative distance between the agent and the obstacles based on the obstacles and the initial position; Step S3: Construct an enclosing cube based on the initial position to obtain the exploration path of the agent, and calculate the coverage rate based on the enclosing cube and the exploration path to obtain the exploration coverage rate; Step S4: Extract features from the obstacle regions to determine the obstacle types of the obstacles, and calculate the influence degree of the obstacles on the agent based on the obstacle types, the exploration coverage rate, and the relative distance; Step S5: Determine the obstacle avoidance strategy based on the influence degree, and adjust the relative distance according to the obstacle avoidance strategy to obtain the distance between the agent and a preset target position, thereby obtaining the target proximity; Step S6: Determine the reward and punishment mechanism based on the target proximity and the influence degree, calculate the exploration time of the agent based on the exploration path and the exploration coverage rate, compare the exploration time with a preset exploration time threshold to obtain a comparison result, and adjust the reward and punishment mechanism according to the comparison result and the target proximity.
2. The intelligent agent navigation method based on dynamic obstacle detection according to claim 1, wherein The process of identifying the obstacle regions in the images based on the RGB images and the depth images, and determining the initial position of the agent includes: Construct a four-layer progressive convolutional network based on the RGB images to obtain multi-dimensional image features; Input the multi-dimensional image features into a pre-trained instance segmentation network to identify and segment each obstacle region in the RGB images; Input the depth images into a pre-trained monocular depth estimation network to obtain corresponding relative depth maps; Use a feature point matching algorithm to extract the matching feature points of the RGB images and the depth images to obtain matching feature points; Determine the initial position of the agent based on the relative depth maps and the matching feature points.
3. The intelligent agent navigation method based on dynamic obstacle detection according to claim 2, characterized in that, The process of determining the relative distance between the agent and the obstacles based on the obstacles and the initial position includes: Determine the three-dimensional positions and sizes of each obstacle based on the obstacles and the relative depth maps to obtain the obstacle positions and obstacle sizes; Calculate the relative distance based on the initial position, the obstacle positions, and the obstacle sizes.
4. The intelligent agent navigation method based on dynamic obstacle detection according to claim 3, wherein The process of Step S3 includes: Construct a three-dimensional coordinate system with the initial position as the center, and construct an enclosing cube based on the three-dimensional coordinate system. The size of the enclosing cube covers the expected exploration area of the agent; Obtain the maximum speed, acceleration, and turning radius of the agent to evaluate the motion ability of the agent; Determine the exploration path based on the motion ability, the enclosing cube, and the distribution information of the obstacles; Divide the enclosing cube into several volume units of equal size. When the agent moves along the exploration path, dynamically mark the volume units passed by the agent. Calculate the exploration coverage rate according to the ratio of the number of volume units passed by the agent to the total number of volume units of the enclosing cube.
5. The intelligent agent navigation method based on dynamic obstacle detection according to claim 4, wherein The process of calculating the exploration coverage rate according to the ratio of the number of volume units passed by the agent to the total number of volume units of the enclosing cube includes: Create a three-dimensional array, or three-dimensional matrix, to record whether each volume unit has been explored by the agent. If the volume unit has been explored, obtain the trajectory point of the agent; When the trajectory point is mapped to a certain volume unit, mark the exploration status of the volume unit as explored; Traverse the three-dimensional array, or the three-dimensional matrix, and count the number of volume units with the exploration status of explored; Calculate the total number of volume units according to the size of the enclosing cube; Obtain the exploration coverage rate according to the ratio of the number of volume units and the total number of volume units.
6. The intelligent agent navigation method based on dynamic obstacle detection according to claim 5, wherein The process of extracting features from the obstacle area to determine the obstacle type of the obstacle includes: Extract features of the shape, color, texture, and depth of the obstacle to obtain multiple feature results; Fuse multiple feature results to obtain a multi-dimensional feature vector; Input the multi-dimensional feature vector into a trained classifier to obtain the obstacle type.
7. The intelligent agent navigation method based on dynamic obstacle detection according to claim 6, wherein The process of calculating the influence degree of the obstacle on the agent according to the obstacle type, the exploration coverage rate, and the relative distance includes: Assign a weight to each obstacle type to obtain an assignment result; Analyze the influence degree of the exploration coverage rate on the obstacle to obtain a first analysis result; Analyze the influence degree of the relative distance on the obstacle to obtain a second analysis result; Calculate the comprehensive influence degree according to the assignment result, the first analysis result, and the second analysis result to obtain the influence degree.
8. The intelligent agent navigation method based on dynamic obstacle detection according to claim 7, wherein The process of step S5 includes: Set thresholds corresponding to different influence degrees to obtain a first degree threshold, a second degree threshold, and a third degree threshold; Compare according to the first degree threshold, the second degree threshold, the third degree threshold, and the influence degree to obtain a corresponding obstacle avoidance strategy; Adjust the relative distance according to the obstacle avoidance strategy to obtain a strategy adjustment distance; Use the Euclidean distance to calculate the relative distance between the strategy adjustment distance and the target distance to obtain the target proximity.
9. The intelligent agent navigation method based on dynamic obstacle detection according to claim 8, characterized in that, The process of determining the reward and punishment mechanism according to the target proximity and the influence degree includes: Set a first target proximity threshold, a second target threshold, and a third target threshold, and compare according to the first target proximity threshold, the second target threshold, the third target threshold, and the target proximity to obtain a comparison result; Assign a score to the first degree threshold, the second degree threshold, and the third degree threshold to obtain a scoring result; Determine the reward and punishment mechanism according to the comparison result and the scoring result.
10. The intelligent agent navigation method based on dynamic obstacle detection according to claim 8, characterized in that, The process of comparing the exploration time with a preset exploration time threshold to obtain a comparison result, and adjusting the reward and punishment mechanism according to the comparison result and the target proximity includes: If the actual exploration time is less than or equal to the exploration time threshold, it is considered that the exploration time is reasonable; If the actual exploration time is greater than the exploration time threshold, it is considered that the exploration time is unreasonable; Among them, when the exploration time is reasonable, when the target proximity is less than or equal to the first target threshold, a high reward is given. When the target proximity is greater than the first target proximity threshold and less than or equal to the second target threshold, a medium reward is given. When the target proximity is greater than the second target proximity threshold, a low reward is given; When the exploration time is unreasonable, if the target proximity is greater than or equal to the second target threshold, a mild penalty is given. When the target proximity is less than the second target threshold, a severe penalty is given.
Citation Information
Patent Citations
Three-dimensional navigation method for network game of mobile terminal
CN106060052A