A Method for Agent Target Search Based on Scene Priors

Through the scene-priority-based agent target search method, the robot can perform map-free navigation and target search in complex environments, solving the problem of traditional navigation methods in the lack of landmarks or GPS signals, and achieving efficient and accurate navigation and target search.

CN115311538BActive Publication Date: 2025-06-27SHANGHAI INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210156851.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-06-27
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

In dynamic, complex and large-scale environments, robot mapping and navigation face challenges, especially in the absence of landmarks or GPS signals, traditional navigation methods are difficult to effectively conduct self-motion estimation and scene information acquisition.

Method used

Using the scene priori-based agent target search method, the robot obtains environmental images, constructs a depth image matrix and a semantic image matrix, performs object relationship feature analysis, obtains spatial semantic point clouds and fusion matrix, generates fusion feature vectors, and trains a value network for target search.

Benefits of technology

Map-free navigation and indoor target search are achieved, navigation accuracy and efficiency are improved, and training time and the risk of network difficulty convergence are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311538B_ABST
    Figure CN115311538B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for intelligent agent target search based on scene prior, which is used for target search of a robot and includes the following steps: confirming target coding information and the target to be searched; obtaining the environmental image of the scene to be searched by the robot, and constructing a depth image matrix and a semantic image matrix according to the environmental image; extracting object relationship feature vectors; constructing a spatial semantic fusion matrix; obtaining semantic map feature vectors according to the spatial semantic map fusion matrix; generating fusion feature vectors according to the object relationship feature vectors, semantic map feature vectors and target coding information; training a value network and a target network according to the fusion feature vectors, and performing target search based on the trained value network after the training is completed. Compared with the prior art, the present invention has the advantages of high navigation accuracy, high search accuracy and efficiency, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of active vision perception, and more particularly to a method for agent target search based on scene prior knowledge. Background Art

[0002] In recent years, the field of robotics research has been dedicated to expanding the capabilities of robots to explore the environment, understand the environment, interact with the environment, and communicate with people. Traditional navigation methods typically use environmental maps for navigation and divide the navigation task into three steps: mapping, localization, and path planning. This method usually requires the prior construction of a 3D map, as well as reliable map localization and path tracking. However, in some cases, artificial landmarks are unknown, or the robot is in an environment where GPS is unavailable, making self-motion estimation or obtaining scene information extremely difficult. For a long time, the problem of robot navigation has been basically solved by a series of distance sensors, such as light detection and ranging, infrared radiation, or sonar navigation and ranging, which are suitable for small-scale static environments (various distance sensors are limited by their respective physical properties). However, in dynamic, complex, and large-scale environments, map building and navigation of robots may face many challenges.

[0003] Recently, the success of data-driven machine learning strategies for various control and perception problems has opened up a new way to overcome the limitations of previous methods. These methods have been widely studied because they do not require map construction, have a lower dependence on the environment, and can perform human-machine interaction. Their key point is to directly learn the mapping between the original observations and the operation tasks in an end-to-end manner. These methods utilize the ability to draw on previous navigation experience in new and similar environments, whether or not there is a map. Reinforcement Learning is commonly used in visual navigation. However, reinforcement learning still has problems such as poor generalization ability, low navigation efficiency, and low accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for agent target search based on scene prior knowledge to overcome the defects of the above-mentioned existing technologies.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A method for agent target search based on scene prior knowledge for target search of a robot, comprising the following steps:

[0007] S1: Confirm the target encoding information and the target to be searched;

[0008] S2: Obtain the environmental image of the scene to be searched by the robot, and construct a depth image matrix and a semantic image matrix according to the environmental image;

[0009] S3: Analyze the object relationship features of the environmental image, identify the objects in the environment, confirm the object with the highest likelihood of relationship with the target to be searched, and extract the object relationship feature vector;

[0010] S4: Obtain the spatial semantic point cloud based on the depth image matrix and the semantic image matrix, and construct a spatial semantic fusion matrix based on the spatial semantic point cloud and the object information in the environment;

[0011] S5: Obtain the semantic map feature vector based on the spatial semantic map fusion matrix;

[0012] S6: Generate a fusion feature vector based on the object relationship feature vector, the semantic map feature vector, and the target encoding information;

[0013] S7: Train the value network and the target network based on the fusion feature vector. After completion of the training, perform target search based on the trained value network.

[0014] Preferably, the step S2 specifically includes:

[0015] S21: Obtain the environmental image of the scene to be searched through the robot. The environmental image includes the RGB image and the depth image of the environment;

[0016] S22: Denote the depth image as the depth image matrix;

[0017] S23: Use the pre-trained semantic segmentation network to calculate the environmental image and generate the semantic image matrix.

[0018] Preferably, the step S3 specifically includes:

[0019] S31: Obtain the scene graph G = {V, E}, where V is the graph node representing different object types in the scene, and E is the graph edge representing the positional relationship between two types of objects. Use the Visual Genome dataset as the source, construct a knowledge graph according to the categories of all objects appearing in the scene to be searched, represent each category as a node in the graph, and use an edge to link between two nodes with an object relationship occurrence frequency greater than 3 in the Visual Genome dataset to generate a graph structure and represent it with a binary adjacency matrix A;

[0020] S32: Construct a graph convolutional neural network. The input is the RGB image of the environmental image, and the output is the spatial relationship feature. Map the spatial relationship feature to 512 dimensions to obtain the object relationship feature vector.

[0021] Preferably, the specific steps of the step S4 include:

[0022] S41: Generate a zero matrix of (C + 2) * (224 * 224), which represents the spatial semantic fusion matrix M. The spatial semantic fusion matrix contains C + 2 layers, where 224 * 224 represents the size of each layer;

[0023] S42: Consider the position and pose P(x t , y t , z t , θ t ) of the robot to generate a spatial point cloud;

[0024] S43: The size of the spatial point cloud is C * W * L * H, where C is the channel of the spatial semantic point cloud, and each channel represents a semantic category. W * L * H are the width, length, and height of the spatial semantic point cloud respectively. Sum in the height dimension and map the three-dimensional point cloud to two dimensions to obtain a two-dimensional mapped feature map of size C * W * L as the first C layers of the spatial semantic fusion matrix;

[0025] S44: Record the path of the robot walking in the C + 1 layer of the spatial semantic fusion matrix, and mark the object with the highest possibility of relationship with the target to be searched in the C + 2 layer.

[0026] Preferably, the way to obtain the spatial point cloud is:

[0027]

[0028] Among them, x, y, and z are the point cloud coordinates respectively, f x , f y is the internal parameter of the camera, c x , c y is the position of the pixel in the semantic image matrix S, D is the depth image matrix, u and v are the pixel coordinates in the semantic image matrix respectively, and R and T are the transformation matrix and rotation matrix of the robot respectively. According to the pose of the robot P(x t , y t , z t , θ t ) to obtain the transformation matrix and rotation matrix of the robot respectively as:

[0029]

[0030]

[0031] Preferably, the step S5 specifically includes:

[0032] S51: Perform normalization processing on the spatial semantic fusion matrix;

[0033] S52: Construct a convolutional neural network, and use the convolutional neural network to process the spatial semantic fusion matrix as input, and output a semantic map feature vector.

[0034] Preferably, the convolutional neural network includes a convolutional layer, a non-linear activation layer, a data normalization layer, a max pooling network, a convolutional layer, a non-linear activation layer, a data normalization layer, a max pooling network, a convolutional layer, a non-linear activation layer, a data normalization layer, a max pooling network, a convolutional layer, and a data normalization layer connected in sequence. Finally, the output of the last data normalization layer is transformed into a one-dimensional vector through matrix transformation, and then through linear transformation, the result of the matrix transformation is converted into a semantic map feature vector.

[0035] Preferably, step S6 specifically includes:

[0036] Concatenate the object relationship feature vector, the semantic map feature vector, and the target encoding information vector to generate a fusion feature vector.

[0037] Preferably, step S7 specifically includes:

[0038] S71: Construct a reward and punishment function:

[0039]

[0040] Among them, R(t,a) is the reward and punishment return, t represents the robot at a certain moment, a represents the action taken by the robot at that moment. When the target category appears in the robot's semantic image matrix S, and it is calculated that the distance between the robot and the target category is less than 0.5m, it means that the robot has found the target;

[0041] S72: Input the fusion feature vector Q into a deep convolutional neural network with initial weights. The machine imitates the navigation strategy of human experts to obtain demonstration experience, and stores the demonstration experience in the initialized experience pool. Then, initialize the value network J with random weight values, initialize the target value network J' as the current value network, and loop through each event to obtain the optimal value network J.

[0042] Preferably, the training process of the value network in S72 is specifically as follows:

[0043] Use the temporal difference method of reinforcement learning to train the value network. Denote the value network J as the current value network, initialize the number of training times to 0, design the experience replay capacity and sampling quantity, set the target network J', initialize the random pose of the robot, set the number of training times, and select an action according to the greedy strategy in the current state:

[0044]

[0045] where a tFor the action to be taken at the next moment, obtain the reward R(t,a) according to the taken action, and the next state S t , store the updated reporting value and state in the experience pool, update the experience pool every preset number of steps, and update the current value network through the gradient descent algorithm until the robot reaches the final state and exceeds the set maximum time t max is 200 actions; otherwise, update the current network to the target network, and obtain the value network J when the training times are reached.

[0046] Compared with the prior art, the present invention has the following advantages:

[0047] 1. The robot target search method of the present invention designs an indoor mapless navigation target search system based on a vision sensor based on the real environment, enabling the robot to search for objects and complete task navigation without the need to build a map, and can achieve mapless target search and indoor navigation tasks.

[0048] 2. In the current research on indoor mapless visual navigation, basically, visual information is used as an input matrix, and reinforcement learning or imitation learning is directly used for training navigation. This method not only has a low navigation success rate and takes a long time, but also some networks are difficult to converge, resulting in training failure. The mapless navigation target search system based on scene prior designed by the present invention first constructs a local semantic map from visual information, greatly improving the training speed and navigation accuracy.

[0049] 3. The present invention adopts scene prior and uses scene prior to explore the environment to find objects, which helps to improve the accuracy and efficiency of target search. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a structural block diagram of the hardware system related to the method of the present invention.

[0051] Figure 2 is the overall flow block diagram of the present invention.

[0052] Figure 3 is a schematic diagram of the scene atlas of the present invention.

[0053] Figure 4 is a schematic diagram of the graph convolutional neural network of the present invention

[0054] Figure 5 is the flowchart of reinforcement learning in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0055] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Note that the following description of the embodiments is only illustrative in nature, and the present invention is not intended to limit the objects to which it applies or its uses, and the present invention is not limited to the following embodiments.

[0056] Embodiment

[0057] A method for intelligent agent target search based on scene prior of the present invention captures RGB images and depth maps of the shooting environment, uses a trained semantic segmentation network to calculate and generate semantic images from the RGB images, and generates a local semantic map based on the depth map and odometer information; records the object relationships that appear multiple times in the Visual Genome data as a prior knowledge matrix, adds object relationships on the basis of the semantic map to generate a spatial relationship semantic map; uses a convolutional network to obtain a spatial semantic map matrix as a data fusion matrix for the local environment; trains a reinforcement learning model as a navigator, takes the data fusion matrix as input, and outputs one of the actions of "forward, left, right, stop" to control the movement direction of the robot. For the target search of the robot, as Figure 1 shown, the equipment adopted by this method mainly consists of a robot equipped with a camera sensor and a lidar and a server. The robot transmits the "seen" information to the server through the camera sensor via WiFi, and converts the visual information on the server. Specifically, as Figure 2 shown, the method includes the following steps:

[0058] S1: Confirm the target encoding information and the target to be searched. Establish a target encoding network and perform encoding to confirm the encoding information of each object in the scene. Specifically, in this embodiment, a human-computer interaction interface is built. Through a text box form of this interaction interface, the user inputs the name of the target to be searched into the text box, and encodes the target after input. The encoder is mainly composed of an LSTM, and each layer contains 128 hidden units.

[0059] S2: Obtain the environmental image of the scene to be searched through the robot, and construct a depth image matrix and a semantic image matrix according to the environmental image.

[0060] Step S2 specifically includes: The robot uses the camera sensor to capture the RGB image and depth map of the environment, which is called the environmental image. The environmental image is a 3*(w*h) image, where w and h are the width and height of the image. The environmental image contains 3 layers, and the size of each layer is (w*h). The depth image is a 1*(w*h) image, and the depth image contains 1 layer, and the size of this layer is (w*h). Denote the depth image as the depth image matrix D. First, use the pre-trained semantic segmentation network Mask-RCNN to process and calculate the environmental image to generate a semantic segmentation result matrix, denoted as the semantic image matrix S.

[0061] S3: Analyze the object relationship features of the environmental image, identify the objects in the environment and confirm the object with the highest possibility of relationship with the target to be searched, and extract the object relationship feature vector.

[0062] As Figure 4 shown, step S3 mainly extracts the relationship information between targets using graph convolutional neural networks (GCNs) based on visual information. The scene prior knowledge is represented in the form of an undirected graph, and the scene graph G = {V, E}, where the nodes in V represent different object categories, and the edge E represents the special positional relationship between two category objects. Use the Visual Genome dataset as the source, and construct a knowledge graph according to the categories of all objects that actually appear in the actual scene in this experiment. Each category is represented as a node in the graph. When the frequency of object relationships in the Visual Genome dataset is greater than 3, an edge is used to link between two nodes to generate a graph structure, which is represented by a binary adjacency matrix A.

[0063] Construct a graph convolutional neural network, with the input being the RGB image. The outputs of each layer of the graph convolutional neural network form a spatial relationship feature matrix Z = [z1, z2,..., z |V| . In this embodiment, there are a total of 83 items in the actual environment, so 83 nodes are designed, and all the nodes are summarized into a feature matrix F = [F1, F2... F |V| . Standardize the matrix A to obtain the matrix

[0064] Therefore, it can be obtained that:

[0065]

[0066] where, H (0) = X, H (β) = Z, where W (α) is the parameter of the α-th layer, and β is the total number of layers of GCNs.

[0067] As Figure 3As shown, the first part of the graph convolutional neural network is a pre-trained RESNet34. The input of this part is an RGB image, and the output is the scores of 1000 classes of objects in ImageNet as an image feature vector. For different nodes, the current image feature vector is mapped to a 512-dimensional feature vector, and then the names of all classes are respectively mapped to 512-dimensional feature vectors using word embeddings. Then the two feature vectors are concatenated to form a 1024-dimensional joint representation for each graph node. The input of the graph convolutional neural network is the adjacency matrix A and the node feature vector. The output of the first two layers is a 1024-dimensional latent feature, and the output of the last layer is the value output for each node, obtaining the feature vector |V|. This feature vector is the semantic encoding information of the current scene and environmental context. Finally, this feature vector is mapped to a 512-dimensional object relationship feature vector e.

[0068] S4: Obtain the spatial semantic point cloud based on the depth image matrix and the semantic image matrix, and construct a spatial semantic fusion matrix according to the spatial semantic point cloud and the object information in the environment.

[0069] Specifically, first generate a all-zero matrix of (C + 2) * (224 * 224), and this matrix represents the spatial semantic fusion matrix M. Among them, the spatial semantic fusion matrix contains C + 2 layers, and v, 224 * 224 represents the size of each layer. The position of the robot is placed in the exact middle of this map, at (112, 112). Considering the position and pose P(x t , y t , z t , θ t ) of the robot, use the following formula to calculate and generate the spatial point cloud:

[0070]

[0071] Among them, x, y, and z are the point cloud coordinates respectively, f x , f y are the internal parameters of the camera, c x , c y are the positions of the pixels in the semantic image matrix S, D is the depth image matrix, u, v are the pixel coordinates in the semantic image matrix respectively, and R, T are the translation matrix and rotation matrix of the robot respectively. According to the pose of the robot being P(x t , y t , z t , θ t ) to obtain the translation matrix and rotation matrix of the robot respectively as:

[0072]

[0073]

[0074] According to the above formula, the spatial semantic point cloud can be calculated through the image feature matrix S and the depth image matrix D. The size of the calculated spatial semantic point cloud is C*W*L*H, where C is the channel of the spatial semantic point cloud, and each channel represents a semantic category. W*L*H are the width, length, and height of the spatial semantic point cloud respectively. Since the computational amount of the spatial semantic point cloud is too large, summation is performed in the high dimension to map the three-dimensional point cloud to two dimensions, obtaining a two-dimensional mapped feature map with the size of C*W*L as the first C layers of the spatial semantic fusion matrix; the path of the robot's movement is recorded in the C+1 layer of the spatial semantic fusion matrix, and the object with the highest possibility of relationship with the target to be searched is marked in the C+2 layer; during the robot's movement, the two-dimensional mapped feature map is added to the spatial semantic fusion matrix M according to the corresponding positions, and the path and the object with the highest possibility of relationship are updated to complete the update of the spatial semantic fusion matrix.

[0075] Furthermore, through the semantic two-dimensional mapped feature map, it can be clearly known which objects exist in the mapped map. Using the relationships between objects in the Visual Genome dataset, the object with the highest possibility of relationship with the target object is statistically calculated and obtained, and the corresponding position of this object is highlighted in the C+2 layer to play a marking role.

[0076] S5: Obtain the semantic map feature vector according to the spatial semantic map fusion matrix;

[0077] Specifically, it includes: S51: Perform normalization processing on the spatial semantic fusion matrix;

[0078] S52: Construct a convolutional neural network, and use the convolutional neural network to process the spatial semantic fusion matrix as the input to output the semantic map feature vector.

[0079] As a way of extracting image features, the convolutional neural network is widely used due to its advantages such as not requiring preprocessing of images and being able to perform additional feature extraction. Moreover, its unique processing method of sharing convolutional kernels can process high-dimensional data. As the number of network layers deepens, the convolutional neural network can extract deep information in images. Therefore, in this step, we use the convolutional neural network to perform feature processing on the images obtained by the camera sensor.

[0080] (1) In this step, the spatial semantic map fusion matrix M is first normalized, and the specific normalization method can be expressed by the following formula:

[0081]

[0082] In the formula, x i * represents the value of the normalized matrix M, x i represents the value of the original matrix M, xmin represents the minimum value of matrix M, x max represents the maximum value of matrix M. Through the above formula, the spatial semantic map fusion matrix M can be normalized.

[0083] (2) Construct a convolutional neural network, which specifically includes the following steps:

[0084] Set the first layer of the convolutional neural network as the convolutional layer. The convolutional kernel of this convolutional layer is a 3*3 matrix, and the number of channels is 64; the input of this convolutional layer is the spatially semantic fusion matrix M after the normalization process in the previous step; the second layer of the convolutional neural network is the non-linear activation layer, and the non-linear activation function is the relu function. The output of the convolutional layer is used as the input of this layer to increase the non-linearity of the network. The third layer of the convolutional neural network is the data normalization layer. The input of this layer is the output of the non-linear activation layer, and the following formula is used to perform normalization calculation on the input:

[0085]

[0086] where, is the output of the normalization layer, x v1 (k) is the output of the non-linear activation layer, k is the channel number, that is, the output of the k-th channel is x v1 (k) , E(x v1 (k) ) is the average value of x v1 (k) and var[x v1 (k) is the variance of x v1 (k) .

[0087] The fourth layer of the convolutional neural network is the max pooling network. The convolutional kernel of the max pooling neural network is a 2*2 matrix. The fifth layer of the convolutional neural network is the convolutional layer. The size of the convolutional kernel of this convolutional layer is a 3*3 matrix, and the number of channels is 64. The input of this convolutional layer is the result output by the max pooling network of the fourth layer of the feature extraction network. The sixth layer of the convolutional neural network is the non-linear activation layer. The non-linear activation function is the relu function. The output of the convolutional layer is used as the input of this layer to increase the non-linearity of the network. The seventh layer of the convolutional neural network is the data normalization layer. The input of this layer is the output of the non-linear activation layer. The eighth layer of the convolutional neural network is the max pooling network. The convolutional kernel of the max pooling neural network is a 2*2 matrix. The ninth layer of the convolutional neural network is the convolutional layer. The size of the convolutional kernel of this convolutional layer is a 3*3 matrix, and the number of channels is 128. The input of this convolutional layer is the result output by the max pooling network. The tenth layer of the convolutional neural network is the non-linear activation layer. The relu function is used as the non-linear activation function. The output of the convolutional layer is used as the input of this layer to increase the non-linearity of the network. The eleventh layer of the convolutional neural network is the data normalization layer. The input of this layer is the output of the non-linear activation layer. The twelfth layer of the convolutional neural network is the max pooling network. The convolutional kernel of the max pooling neural network is a 2*2 matrix. The thirteenth layer of the convolutional neural network is the convolutional layer. The convolutional kernel of this convolutional layer is a 3*3 matrix, and the number of channels is 512. The input of this convolutional layer is the result output by the max pooling network. The fourteenth layer of the convolutional neural network is the data normalization layer. The input is the output result of the thirteenth layer. Through matrix transformation, the output of the data normalization layer becomes a one-dimensional vector. Using linear transformation, the matrix transformation becomes a semantic map feature vector f of 1*1*128.

[0088] S6: Generate a fused feature vector based on the object relationship feature vector, the semantic map feature vector, and the target encoding information. Specifically, in this embodiment, the feature vectors e, f, and the target encoding information are concatenated to generate a fused feature vector Q.

[0089] Step S7 uses the deep convolutional neural network in deep learning and the temporal difference method in reinforcement learning to train the value network model to achieve the target search and navigation of the robot, specifically including:

[0090] S71: Construct a reward and punishment function:

[0091]

[0092] Among them, R(t,a) is the reward and punishment return. t represents the robot at a certain moment, and a represents the action taken by the robot at this moment. When the target category appears in the robot semantic image matrix S and the calculated distance between the robot and the target category is less than 0.5m, it means that the robot has found the target;

[0093] S72: Input the fused feature vector Q into a deep convolutional neural network with initial weights. The machine imitates the navigation strategy of human experts to obtain demonstration experience and stores the demonstration experience in the initialized experience pool. Initialize the value network J with random weight values, initialize the target value network J' as the current value network, and loop through each event to obtain the optimal value network J.

[0094] In this embodiment, as Figure 5 shown, the training process of the value network in S72 is specifically as follows:

[0095] Use the temporal difference method of reinforcement learning to train the value network. Denote the value network J as the current value network, initialize the number of training times to 0, design the experience replay capacity to be 50000, the sampling quantity to be 200, set the target network J', initialize the random pose of the robot, the number of training times to 1000000, and select an action according to the greedy strategy in the current state:

[0096]

[0097] where a t is the action taken at the next moment. At the beginning, the robot doesn't know how to take an action, so it can only be random. However, when there is a certain amount of experience, the robot will look for the action that can obtain the maximum reward to execute. Obtain the reward R(t,a) and the next state S t according to the taken action, store the updated reported value and state in the experience pool, update the experience pool every 3000 steps, and update the current value network through the gradient descent algorithm until the robot reaches the final state and exceeds the set maximum time t max which is 200 actions. Otherwise, update the current network to the target network, and obtain the value network J when the training times are reached.

[0098] The above embodiments are only examples and do not represent limitations on the scope of the present invention. These embodiments can also be implemented in various other ways and can be subject to various omissions, substitutions, and changes without departing from the technical idea of the present invention.

Claims

1. A method for intelligent agent target search based on scene prior, characterized in that, Target search for a robot, including the following steps: S1: Confirm the target encoding information and the target to be searched; S2: Obtain the environmental image of the scene to be searched by the robot, and construct a depth image matrix and a semantic image matrix based on the environmental image; S3: Conduct an object relationship feature analysis on the environmental image, identify the objects in the environment and confirm the object with the highest possibility of relationship with the target to be searched, and extract the object relationship feature vector; S4: Obtain the spatial semantic point cloud based on the depth image matrix and the semantic image matrix, and construct a spatial semantic fusion matrix based on the spatial semantic point cloud and the object information in the environment; S5: Obtain the semantic map feature vector based on the spatial semantic fusion matrix; S6: Generate a fusion feature vector based on the object relationship feature vector, the semantic map feature vector and the target encoding information; S7: Construct a value network and a target network for target search, train the value network and the target network according to the fusion feature vector, and conduct target search based on the trained value network after completion of training; Step S3 specifically includes: S31: Obtain the scene graph G = {V, E}, where V is the graph node representing different object types in the scene, and E is the graph edge representing the positional relationship between two types of objects. Use the Visual Genome dataset as the source, construct a knowledge graph according to the categories of all objects appearing in the scene to be searched, represent each category as a node in the graph, and use an edge to link between two nodes with an object relationship occurrence frequency greater than 3 in the Visual Genome dataset to generate a graph structure and represent it with a binary adjacency matrix A; S32: Construct a graph convolutional neural network, with the input being the RGB image of the environmental image and the output being the spatial relationship feature, and map the spatial relationship feature to 512 dimensions to obtain the object relationship feature vector; The specific steps of step S4 include: S41: Generate a (C + 2) * (224 * 224) all-zero matrix, which represents the spatial semantic fusion matrix M. The spatial semantic fusion matrix contains C + 2 layers, where 224 * 224 represents the size of each layer; S42: Consider the position and orientation P(x t , y t , z t , θ t ) to generate a spatial point cloud; S43: The size of the spatial point cloud is C * W * L * H, where C is the channel of the spatial semantic point cloud, and each channel represents a semantic category. W * L * H are respectively the width, length and height of the spatial semantic point cloud. Sum in the height dimension and map the three-dimensional point cloud to two dimensions to obtain a two-dimensional mapping feature map with a size of C * W * L as the first C layers of the spatial semantic fusion matrix; S44: Record the path of the robot in the C + 1 layer of the spatial semantic fusion matrix, and mark the object with the highest possibility of relationship with the target to be searched in the C + 2 layer; S45: Obtain the latest spatial point cloud, path and the object with the highest possibility of relationship with the target to be searched of the robot in real time, and update the spatial semantic fusion matrix.

2. The method for intelligent agent target search based on scene prior according to claim 1, characterized in that Step S2 specifically includes: S21: Obtain the environmental image of the scene to be searched by the robot. The environmental image includes the RGB image and the depth image of the environment; S22: Denote the depth image as the depth image matrix; S23: Use the pre-trained semantic segmentation network to calculate the environmental image and generate the semantic image matrix.

3. The method for intelligent agent target search based on scene prior according to claim 1, wherein, The acquisition method of the spatial point cloud is as follows: Among them, x, y, and z are the point cloud coordinates, and f x , f y are the internal parameters of the camera, c x , c y is the position of the pixel in the semantic image matrix S, D is the depth image matrix, u and v are the pixel coordinates in the semantic image matrix, and R and T are the translation matrix and rotation matrix of the robot respectively. According to the pose of the robot being P(x t , y t , z t , θ t ), the translation matrix and rotation matrix of the robot are obtained as follows:

4. The method for an agent's target search based on scene prior according to claim 1, wherein Step S5 specifically includes: S51: Normalize the spatial semantic fusion matrix; S52: Construct a convolutional neural network, and use the convolutional neural network to process the spatial semantic fusion matrix as input, and output a semantic map feature vector.

5. The method for intelligent agent target search based on scene prior according to claim 4, wherein The convolutional neural network includes a convolutional layer, a non-linear activation layer, a data normalization layer, a max pooling network, a convolutional layer, a non-linear activation layer, a data normalization layer, a max pooling network, a convolutional layer, a non-linear activation layer, a data normalization layer, a max pooling network, a convolutional layer, and a data normalization layer connected in sequence. Finally, the output of the last data normalization layer is transformed into a one-dimensional vector through matrix transformation, and then through linear transformation, the result of the matrix transformation is converted into a semantic map feature vector.

6. A method for intelligent agent target search based on scene prior according to claim 1, characterized in that, Step S6 specifically includes: Concatenate the object relationship feature vector, the semantic map feature vector, and the target encoding information vector to generate a fused feature vector.

7. A method for intelligent agent target search based on scene prior according to claim 1, characterized in that Step S7 specifically includes: S71: Construct a reward and punishment function: Among them, R(t,a) is the reward and punishment return, t represents the robot at a certain moment, a represents the action taken by the robot at that moment. When the target category appears in the robot semantic image matrix S, and the calculated distance between the robot and the target category is less than 0.5m, it means that the robot has found the target; S72: Input the fused feature vector Q into a deep convolutional neural network with initial weights. The machine imitates the navigation strategy of human experts to obtain demonstration experience, and stores the demonstration experience in the initialized experience pool. Then initialize the value network J with random weight values, initialize the target network J’ as the current value network, and loop through each event to obtain the optimal value network J.

8. A method for intelligent agent target search based on scene prior according to claim 7, characterized in that, The training process of the value network in S72 is specifically as follows: Use the temporal difference method of reinforcement learning to train the value network. Denote the value network J as the current value network, initialize the number of training times to 0, design the experience replay capacity and the sampling quantity, set the target network J’, initialize the random pose of the robot, set the number of training times, and select an action according to the greedy strategy in the current state: where a t is the action taken at the next moment. The reward R(t, a) is obtained based on the action taken, and the next state S t . The updated reported value and state are stored in the experience pool. The experience pool is updated every preset number of steps. The current value network is updated through the gradient descent algorithm until the robot reaches the final state and exceeds the set maximum time t max is 200 actions; otherwise, the current network is updated to the target network, and the value network J is obtained when the number of training times is reached.