A trajectory prediction method, device and equipment for a robot to inspect a target and a medium
By combining a variational adaptive transformer with a multi-head self-attention and adaptive graph convolutional network module to create a trajectory prediction model, the problem of insufficient reliability of robot trajectory prediction in complex scenarios in existing technologies is solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202310028307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-01-09
AI Technical Summary
In existing technologies, robot trajectory prediction methods based on deterministic generative models have difficulty effectively taking into account the randomness and complexity of multidimensional time series, resulting in reduced reliability of trajectory prediction in complex scenarios.
A variational adaptive transformer is used, combined with a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module, to construct a trajectory prediction model. By capturing time dependence and integrating features, the randomness and complexity of motion trajectories are taken into account.
It improves the accuracy and reliability of robot trajectory prediction, enabling effective target trajectory prediction in complex scenarios.
Smart Images

Figure CN116645395B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a method, apparatus, equipment and medium for predicting the trajectory of a robot inspection target. Background Technology
[0002] In the current construction of various smart scenarios, intelligent transportation systems have become an active research area because of their potential to improve system efficiency and decision-making. Intelligent driving technology needs to be integrated into the research field of outdoor security robots to enhance the safety performance of robot inspection and other activities. Specifically, trajectory prediction of targets can increase the robot's reaction time in complex scenarios, thereby improving its driving safety. Therefore, trajectory prediction for robots is a crucial issue in the development of robotics technology.
[0003] Currently, when predicting robot trajectories, the modeling and analysis of multidimensional time series based on transformers typically employs deterministic generative models. This neglects the randomness within the multidimensional time series, making it difficult for existing transformers to model multidimensional time series with complex distributions. Consequently, the reliability of the obtained trajectory prediction model is reduced, hindering the ability to predict robot trajectories in complex scenarios. Summary of the Invention
[0004] To address the aforementioned technical problems, this specification provides one or more embodiments of a method, apparatus, device, and medium for predicting the trajectory of a robot inspection target.
[0005] One or more embodiments of this specification employ the following technical solutions:
[0006] This specification provides one or more embodiments of a method for predicting the trajectory of a target to be inspected by a robot, the method comprising:
[0007] The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices.
[0008] The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected.
[0009] The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0010] Optionally, in one or more embodiments of this specification, before inputting the trajectory data sequence into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, the method further includes:
[0011] The components of the initial variational adaptive transformer for the trajectory prediction model are determined based on the converter framework;
[0012] The encoder module for obtaining the initial variational adaptive transformer is based on the integration of the preset multi-head self-attention module and the preset adaptive graph convolutional network module.
[0013] The encoder module acquires the time of the preset training sample sequence and the correlation information of each channel in the encoder, so as to output the dynamic hidden state of the training sample sequence based on the correlation information.
[0014] Based on the dynamic hidden state, the decoder module of the initial variational adaptive transformer is reconstructed;
[0015] The encoder module and the decoder module are combined based on the positions of the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer.
[0016] Optionally, in one or more embodiments of this specification, before the encoder module that integrates the preset multi-head self-attention module and the preset adaptive graph convolutional network module to obtain the initial variational adaptive transform, the method further includes:
[0017] Obtain the preset number of attention layers in the initial multi-head self-attention module, and input the attention mechanism calculation formula based on the feature matrix corresponding to each preset number of attention layers to obtain the output of each predicted attention layer in the initial multi-head self-attention module;
[0018] The output layers of each of the predicted attention layers are connected through fully connected layers to construct the multi-head self-attention module;
[0019] The number of channels of the encoder is determined based on the multi-head self-attention module, and the number of channels embedded with the input of the multi-head self-attention module is set as the number of outputs;
[0020] The adaptive graph convolutional network module is obtained by reconstructing the initial adaptive graph convolutional network based on the number of inputs to the multi-head self-attention module embedded in each channel as outputs.
[0021] Optionally, in one or more embodiments of this specification, the decoder module that reconstructs the initial variational adaptive transformer based on the dynamic hidden state specifically includes:
[0022] Obtain the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence;
[0023] The dynamic hidden state is calculated based on a preset nonlinear activation function to obtain the mean of the hidden state variable of the Gaussian distribution corresponding to each time step of the training sample sequence.
[0024] The hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence are determined based on the mean of the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence and the covariance parameter of the Gaussian distribution.
[0025] The hidden state variable is used as the input to the probability generation module, and the dynamic hidden state is added to the probability generation module to reconstruct the probability generation module. The reconstructed probability generation module is then used as the decoder module of the initial variational adaptive transformer.
[0026] Optionally, in one or more embodiments of this specification, after combining the encoder module and the decoder module based on the positions corresponding to the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer, the method further includes:
[0027] The loss function of the prediction model is determined by obtaining the prediction loss function of the preset initial trajectory prediction model and the reconstruction loss function of the preset variational adaptive transformer.
[0028] The loss function is calculated based on the preset marginal likelihood definition formula to obtain the initial trajectory prediction model based on the preset variational adaptive transformer;
[0029] Historical trajectory data sequences of each inspection target in the inspection scenario of the target robot are collected as training samples to train the initial prediction model based on the training samples, thereby obtaining the trajectory prediction model obtained by training the preset variational adaptive transformer.
[0030] Optionally, in one or more embodiments of this specification, the step of acquiring image data of the target to be inspected based on one or more acquisition devices pre-installed on the target robot specifically includes:
[0031] Obtain the device type of the target robot and determine the inspection scenario of the target robot;
[0032] Historical inspection records corresponding to the device type and the inspection scenario are obtained based on the preset management system, and the image acquisition frequency of the target robot is determined based on the historical inspection records.
[0033] Based on the image acquisition frequency, the target robot is pre-equipped with one or more acquisition devices to acquire images of the target to be inspected, thereby obtaining image data of the target to be inspected.
[0034] Optionally, in one or more embodiments of this specification, the step of obtaining the position coordinates of the target to be inspected based on the image data, and determining the trajectory data sequence of the target to be inspected according to the acquisition time corresponding to the position coordinates of the target to be inspected, specifically includes:
[0035] The acquisition time of each image data is obtained, and the location coordinates of the target to be inspected corresponding to each image data are obtained;
[0036] The acquisition time corresponding to each of the coordinate positions is determined, and the position coordinates of the target to be inspected are sorted based on the acquisition time to obtain the trajectory data sequence of the target to be inspected.
[0037] This specification provides one or more embodiments of a trajectory prediction device for a robot inspection target, the device comprising:
[0038] The acquisition unit is used to acquire image data of the target to be inspected based on one or more acquisition devices pre-installed on the target robot.
[0039] The determining unit is used to obtain the position coordinates of the target to be inspected based on the image data, and to determine the trajectory data sequence of the target to be inspected according to the acquisition time corresponding to the position coordinates of the target to be inspected;
[0040] The prediction unit is used to input the trajectory data sequence into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0041] This specification provides one or more embodiments of a trajectory prediction device for a robot inspection target, comprising:
[0042] At least one processor; and,
[0043] A memory communicatively connected to the at least one processor; wherein,
[0044] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:
[0045] The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices.
[0046] The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected.
[0047] The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0048] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:
[0049] The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices.
[0050] The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected.
[0051] The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0052] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0053] The trajectory data sequence is input into the trajectory prediction model trained according to a pre-set variational adaptive transformer for prediction. The multi-head self-attention block in the trajectory prediction model is used to capture temporal dependencies. The feature integration is achieved based on the adaptive convolutional network module, and the probability generation module is used as the encoder of the network structure. This fully considers the temporal dependencies and randomness of the motion trajectory of the target to be inspected, improves the accuracy of prediction, and makes this prediction method applicable to complex scenarios. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0055] Figure 1 This is a schematic flowchart of a method for predicting the trajectory of a robot inspection target, provided in an embodiment of this specification.
[0056] Figure 2 This is a schematic diagram of the generation structure of a trajectory prediction model in a certain application scenario provided in the embodiments of this specification;
[0057] Figure 3 This is a schematic diagram of the internal structure of a trajectory prediction device for a robot inspection target provided in an embodiment of this specification;
[0058] Figure 4 This is a schematic diagram of the internal structure of a trajectory prediction device for robot inspection targets provided in an embodiment of this specification;
[0059] Figure 5 This is a schematic diagram of the internal structure of a non-volatile storage medium provided in the embodiments of this specification. Detailed Implementation
[0060] This specification provides a method, apparatus, device, and medium for predicting the trajectory of a target being inspected by a robot.
[0061] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0062] like Figure 1 As shown, this specification provides a flowchart illustrating a method for predicting the trajectory of a target being inspected by a robot, based on one or more embodiments. Figure 1 It can be seen that the method includes the following steps:
[0063] S101: Acquire image data of the target to be inspected based on one or more acquisition devices pre-installed on the target robot.
[0064] With the development of robotics technology, the use of inspection robots is becoming increasingly frequent in various application fields. To enhance safety during robot inspections, trajectory prediction of the target to be inspected can be used to enable the robot's reaction in complex scenarios, providing more time for the robot to react and thus improving safety during operation. To achieve trajectory prediction of the target to be inspected, the embodiments in this specification first use one or more pre-installed acquisition devices on the target robot to acquire images of the target, obtaining image data of the target. It should be noted that the acquisition device can be a camera.
[0065] Specifically, in one or more embodiments of this specification, acquiring image data of the target to be inspected based on one or more pre-installed acquisition devices on the target robot includes the following steps:
[0066] First, the device type of the target robot is obtained, and the inspection scenario of the target robot is determined, such as an inspection robot in a warehouse scenario or an inspection robot in a vehicle tunnel. Since the inspection tasks corresponding to different types of robots in different inspection scenarios are different, the requirements for the degree of trajectory control of the inspection target also vary. Therefore, after obtaining the device type and inspection scenario of the target robot, historical inspection records corresponding to the device type and inspection scenario are obtained according to the robot's pre-set management system. Based on these historical inspection records, the image acquisition frequency of the target robot is determined. Then, according to the image acquisition frequency, one or more acquisition devices pre-installed on the target robot are controlled to perform image acquisition of the target to be inspected, obtaining image data of the target.
[0067] S102: Obtain the position coordinates of the target to be inspected based on the image data, and determine the trajectory data sequence of the target to be inspected according to the acquisition time corresponding to the position coordinates of the target to be inspected.
[0068] After obtaining the image data based on the above steps, in order to predict the trajectory of the target to be inspected by the robot at the next moment, it is necessary to obtain the trajectory point record of the target to be inspected. Therefore, in this embodiment of the specification, the position coordinates of the target to be inspected are obtained based on the image data, and the image data of the target to be inspected is determined based on the acquisition time corresponding to the position coordinates of the target to be inspected.
[0069] Specifically, in one or more embodiments of this specification, the location coordinates of the target to be inspected are obtained based on image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected. The specific process includes the following steps: First, the acquisition time of each image data is obtained, and the location coordinates of the target to be inspected corresponding to each image data are obtained. Then, the acquisition time corresponding to each coordinate position is determined, and the location coordinates of the target to be inspected are sorted according to the acquisition time to obtain the trajectory data sequence of the target to be inspected.
[0070] S103: Input the trajectory data sequence into the trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0071] Robots may face complex scenarios during the inspection of targets, and the movement trajectory of these targets may exhibit randomness. Existing methods generally rely on deterministic generation models for prediction, which makes it difficult for existing prediction models to fully consider the randomness and complexity of the target's multi-dimensional time-series motion trajectory when modeling it. Therefore, to address this issue, in this embodiment, the trajectory data sequence is input into a trajectory prediction model trained using a pre-set variational adaptive transformer to obtain the trajectory prediction result of the inspection target. It should be noted that the pre-set variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0072] Furthermore, in one or more embodiments of this specification, before inputting the trajectory data sequence into the trajectory prediction model trained based on a preset variational adaptive transformer for prediction, the method further includes the following process: First, the components of the initial variational adaptive transformer of the trajectory prediction model are determined based on the transformer framework, thereby obtaining the connection relationship of each component. Then, the encoder module of the initial variational adaptive transformer is obtained by integrating the preset multi-head self-attention module and the preset adaptive graph convolutional network module. The encoder module obtains the correlation information between the time of the preset training sample sequence and each channel in the encoder, and outputs the dynamic hidden state of the training sample sequence based on the correlation information. Based on the dynamic hidden state obtained in the above process, the decoder module of the initial variational adaptive transformer is reconstructed. Then, the encoder module and the decoder module are combined according to the positions corresponding to the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer. This process fully considers the randomness within the motion trajectory sequence of the inspection robot, uses the probability generation module as the decoder and the transformer as the encoder, and incorporates the dynamic hidden state into the generation process, fully considering the temporal dependence and randomness of the motion trajectory sequence, thus improving the reliability of the prediction.
[0073] Specifically, in one or more embodiments of this specification, before the encoder module for obtaining the initial variational adaptive transform based on the integration of a preset multi-head self-attention module and a preset adaptive graph convolutional network module, the method further includes the following steps:
[0074] First, the preset number of attention layers in the initial multi-head self-attention module is obtained. Then, based on the feature matrices corresponding to each preset attention layer, the attention mechanism calculation formula is input to obtain the output of each predicted attention layer in the initial multi-head self-attention module. Next, the output layers of each predicted attention layer are connected through fully connected layers to construct the multi-head self-attention module. For example, in a certain application scenario, a multi-head self-attention module (MSA) is introduced to capture the temporal dependency of the motion trajectory of the target to be inspected. The output of this multi-head self-attention module is:
[0075]
[0076] Among them, the multidimensional time series, which is the trajectory data time series of the target to be detected, is x. n ={x 1,n ,x 2,n ,…x T,n}, n=1…N is the number of time series of trajectory data, and T is the x n The duration of X∈R. V×T The input is represented by con(·), which represents the concatenation operation. Where i∈{1,2,…,m} and m are the number of self-attention layers, Q i K i V i These represent the query, key, and value in different self-attention layer matrix forms, respectively; d k Represents the dimension of the vector key. Assign V i =X to preserve the meaning of channels in the MTS in order to capture their subsequent relationships. It can be understood that O∈R (V×T×m) This is the output of the multi-head self-attention module. Then, the number of encoder channels is determined based on the multi-head self-attention module, and the number of inputs embedded in the multi-head self-attention module for each channel is set as the output. Thus, based on the number of inputs embedded in the multi-head self-attention module for each channel as outputs, the initial adaptive graph convolutional network is reconstructed to obtain the adaptive graph convolutional network module. For example, in a certain application scenario, to avoid the problem that traditional converters using a feedforward network after the MSA module cannot effectively capture the different relationships between different channels, an adaptive graph convolutional network module is introduced. Specifically, using the outputs from the multi-head self-attention module and the channel embedding α, the adaptive graph convolutional network module automatically discovers channel dependencies as shown below:
[0077]
[0078] Where A∈R V×V The relation matrix is calculated using a symmetric similarity function, and α is updated to adapt to the multidimensional time series data. Conv(·) represents the convolution operation, used to summarize the multi-head information output by the multi-head self-attention module into H∈R. V×T . It is a normalized symmetric adjacency matrix. W∈R K′×V It is a GCN filter. It combines multi-dimensional time series and channel relationship information into... Subsequently, multi-head self-attention and feedforward network blocks are further applied to explore the representation of multidimensional time series and obtain the dynamic hidden state h∈R. K×T Then, you can choose between a reconstruction decoder or a prediction decoder for different tasks.
[0079] Specifically, in one or more embodiments of this specification, the decoder module for reconstructing the initial variational adaptive transformer based on the dynamic hidden state includes the following process:
[0080] The hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence are obtained. Then, the dynamic hidden state is calculated based on a preset nonlinear activation function to obtain the mean of the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence. Based on the mean of the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence and the covariance parameter corresponding to the Gaussian distribution, the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence are determined. The hidden state variables are used as input to the probability generation module, and the dynamic hidden state is added to the probability generation module to reconstruct the probability generation module. The reconstructed probability generation module is then used as the decoder module of the initial variational adaptive transformer. Specifically, in a certain application scenario of this specification...
[0081] To account for the randomness within the multidimensional time series, a Gaussian-distributed latent variable z is first introduced at each time step. t,n ∈R K Then the generation process is defined as:
[0082]
[0083] Where, μ t,n and σ t,n It is z t,n The mean and covariance parameters. h t-1,n ∈R K′ Let z represent the deterministic hidden state of the probability generation module. t,n and h t,n This is incorporated into the generation process to account for temporal dependencies and randomness. f(·) refers to the nonlinear activation function. The resulting generative model is the probabilistic generation module. Compared with the generation modules of previous multidimensional time series model methods based on converters, the probabilistic generation module in this specification uses the implicit semantic structure of each channel as a probabilistic embedding vector and captures the relationship between them based on the similarity of the channel embeddings, while taking into account the nondeterminism, thus improving reliability.
[0084] Furthermore, in a certain application scenario of the embodiments of this specification, the preset variational adaptive transformer also includes an inference module. The purpose of the inference module is to transform the input x... n Mapping to z n To account for local time dependencies, we first consider x... t,n Apply convolution operation as Then, using a multi-head self-attention module and an adaptive graph convolutional network module, we summarize the time and channel information of the input multidimensional time series as follows:
[0085]
[0086] Given a time series Apply linear projection and combine it with location embedding to obtain:
[0087]
[0088] by As input, we then apply multi-head self-attention and feedforward network blocks to obtain the dynamic hidden state h. t Based on the variational autoencoder model, a variational distribution of a Gaussian distribution is defined. To approximate the true posterior distribution p(z) t,n |-), and will the dynamic hierarchical h t The parameters mapped to them are:
[0089] μ t,n =f(C xμ h t,n +b x,μ ), σ t,n =softplus(f(C x,σ h t,n +b xσ (7)
[0090] Where C x,μ C x,σ ∈R K×V ,b x,μ ,b x,σ ∈R K These are all learnable parameters of the inference network.
[0091] Furthermore, in one or more embodiments of this specification, after combining the encoder module and the decoder module based on the positions corresponding to the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer, the method further includes the following steps:
[0092] First, the prediction loss function of the pre-set initial trajectory prediction model and the reconstruction loss function of the pre-set variational adaptive transformer are obtained to determine the loss function of the prediction model. Then, the overall loss function is calculated according to the pre-defined marginal likelihood formula to obtain the initial trajectory prediction model based on the pre-set variational adaptive transformer. That is, in a certain application scenario of this specification, the prediction-based model is good at capturing the periodic information of multidimensional time series, while the reconstruction-based model can explore the global distribution of the time series. In order to combine the complementary advantages of the two models to improve the representation capability of multidimensional time series, the optimization function is designed as a combination of prediction and reconstruction losses, and the marginal likelihood is defined as follows:
[0093]
[0094] It should be noted that the first and second terms in the above equation are the losses for reconstruction and prediction, respectively. Then, historical trajectory data sequences of each inspection target in the target robot's inspection scenario are collected as training samples. Based on these training samples, the initial prediction model is trained to obtain the trajectory prediction model trained by the preset variational adaptive transformer.
[0095] like Figure 2 The diagram shown illustrates the generation structure of a trajectory prediction model for a specific application scenario, as provided in one or more embodiments of this specification. Figure 2 As can be seen, in the embodiments of this specification, a transformer framework is introduced into the temporal data modeling process as an encoder module. Specifically, as mentioned above, a multi-head self-attention block is introduced to capture temporal dependencies; an adaptive convolutional network module is introduced to integrate features. A probability generation module is introduced as the encoder of the network structure, where the output features of the transformer at each time step are introduced into the generation process. Temporal correlations are considered in the generation process, and the nondeterminism is taken into account through probabilistic modeling. Then, based on this trajectory prediction model, the input x is... n Mapping to z n Perform reasoning and apply the above formula. The trajectory prediction model is trained and tested using the optimized model. The learned probabilistic latent representation is then input into the predictor to make predictions at multiple time points, and the prediction results are obtained.
[0096] like Figure 3 As shown, this specification provides a schematic diagram of the internal structure of a trajectory prediction device for robot inspection targets in one or more embodiments. Figure 3 It is known that a trajectory prediction device for robot inspection targets includes:
[0097] The acquisition unit is used to acquire image data of the target to be inspected based on one or more acquisition devices pre-installed on the target robot.
[0098] The determining unit is used to obtain the position coordinates of the target to be inspected based on the image data, and to determine the trajectory data sequence of the target to be inspected according to the acquisition time corresponding to the position coordinates of the target to be inspected;
[0099] The prediction unit is used to input the trajectory data sequence into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0100] like Figure 4As shown, one or more embodiments of this specification provide a trajectory prediction device for robot inspection targets, the device comprising:
[0101] At least one processor; and,
[0102] A memory communicatively connected to the at least one processor; wherein,
[0103] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:
[0104] The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices.
[0105] The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected.
[0106] The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0107] like Figure 5 As shown, this specification provides a schematic diagram of the internal structure of a non-volatile storage medium according to one or more embodiments. Figure 5 It is known that a non-volatile storage medium stores computer-executable instructions, the computer-executable instructions including:
[0108] The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices.
[0109] The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected.
[0110] The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module.
[0111] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0112] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0113] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A trajectory prediction method for a robot-inspected target, characterized in that, The method includes: The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices. The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected. The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module; Before inputting the trajectory data sequence into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, the method further includes: The components of the initial variational adaptive transformer for the trajectory prediction model are determined based on the converter framework; The encoder module for obtaining the initial variational adaptive transformer is based on the integration of the preset multi-head self-attention module and the preset adaptive graph convolutional network module. The encoder module acquires the time of the preset training sample sequence and the correlation information of each channel in the encoder, so as to output the dynamic hidden state of the training sample sequence based on the correlation information. Based on the dynamic hidden state, the decoder module of the initial variational adaptive transformer is reconstructed; wherein, the decoder module is a probabilistic generative model; The encoder module and the decoder module are combined based on the positions of the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer.
2. The trajectory prediction method for a robot inspection target according to claim 1, characterized in that, Before the encoder module that integrates the preset multi-head self-attention module and the preset adaptive graph convolutional network module to obtain the initial variational adaptive transform, the method further includes: Obtain the preset number of attention layers in the initial multi-head self-attention module, and input the attention mechanism calculation formula based on the feature matrix corresponding to each preset number of attention layers to obtain the output of each predicted attention layer in the initial multi-head self-attention module; The output layers of each of the predicted attention layers are connected through fully connected layers to construct the multi-head self-attention module; The number of channels of the encoder is determined based on the multi-head self-attention module, and the number of channels embedded with the input of the multi-head self-attention module is set as the number of outputs; The adaptive graph convolutional network module is obtained by reconstructing the initial adaptive graph convolutional network based on the number of inputs to the multi-head self-attention module embedded in each channel as outputs.
3. The trajectory prediction method for a robot inspection target according to claim 1, characterized in that, The decoder module that reconstructs the initial variational adaptive transformer based on the dynamic hidden state specifically includes: Obtain the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence; The dynamic hidden state is calculated based on a preset nonlinear activation function to obtain the mean of the hidden state variable of the Gaussian distribution corresponding to each time step of the training sample sequence. The hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence are determined based on the mean of the hidden state variables of the Gaussian distribution corresponding to each time step of the training sample sequence and the covariance parameter of the Gaussian distribution. The hidden state variable is used as the input to the probability generation module, and the dynamic hidden state is added to the probability generation module to reconstruct the probability generation module. The reconstructed probability generation module is then used as the decoder module of the initial variational adaptive transformer.
4. The trajectory prediction method for a robot inspection target according to claim 1, characterized in that, After combining the encoder module and the decoder module based on the positions of the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer, the method further includes: The loss function of the prediction model is determined by obtaining the prediction loss function of the preset initial trajectory prediction model and the reconstruction loss function of the preset variational adaptive transformer. The loss function is calculated based on the preset marginal likelihood definition formula to obtain the initial trajectory prediction model based on the preset variational adaptive transformer; Historical trajectory data sequences of each inspection target in the inspection scenario of the target robot are collected as training samples. The initial trajectory prediction model is trained based on the training samples to obtain the trajectory prediction model obtained by training the preset variational adaptive transformer.
5. The trajectory prediction method for a robot inspection target according to claim 1, characterized in that, The step of acquiring image data of the target to be inspected based on one or more pre-installed acquisition devices on the target robot specifically includes: Obtain the device type of the target robot and determine the inspection scenario of the target robot; Historical inspection records corresponding to the device type and the inspection scenario are obtained based on the preset management system, and the image acquisition frequency of the target robot is determined based on the historical inspection records. Based on the image acquisition frequency, the target robot is pre-equipped with one or more acquisition devices to acquire images of the target to be inspected, thereby obtaining image data of the target to be inspected.
6. The trajectory prediction method for a robot inspection target according to claim 1, characterized in that, The step of obtaining the location coordinates of the target to be inspected based on the image data, and determining the trajectory data sequence of the target to be inspected according to the acquisition time corresponding to the location coordinates of the target to be inspected, specifically includes: The acquisition time of each image data is obtained, and the location coordinates of the target to be inspected corresponding to each image data are obtained; The acquisition time corresponding to each of the aforementioned location coordinates is determined, and the location coordinates of the target to be inspected are sorted based on the acquisition time to obtain the trajectory data sequence of the target to be inspected.
7. A trajectory prediction device for robot inspection targets, characterized in that, The device includes: The acquisition unit is used to acquire image data of the target to be inspected based on one or more acquisition devices pre-installed on the target robot. The determining unit is used to obtain the position coordinates of the target to be inspected based on the image data, and to determine the trajectory data sequence of the target to be inspected according to the acquisition time corresponding to the position coordinates of the target to be inspected; The prediction unit is used to input the trajectory data sequence into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module. Before inputting the trajectory data sequence into the trajectory prediction model trained based on a preset variational adaptive transformer for prediction, the following steps are also included: The components of the initial variational adaptive transformer for the trajectory prediction model are determined based on the converter framework; The encoder module for obtaining the initial variational adaptive transformer is based on the integration of the preset multi-head self-attention module and the preset adaptive graph convolutional network module. The encoder module acquires the time of the preset training sample sequence and the correlation information of each channel in the encoder, so as to output the dynamic hidden state of the training sample sequence based on the correlation information. Based on the dynamic hidden state, the decoder module of the initial variational adaptive transformer is reconstructed; wherein, the decoder module is a probabilistic generative model; The encoder module and the decoder module are combined based on the positions of the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer.
8. A trajectory prediction device for robot inspection targets, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices. The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected. The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module. Before inputting the trajectory data sequence into the trajectory prediction model trained based on a preset variational adaptive transformer for prediction, the following steps are also included: The components of the initial variational adaptive transformer for the trajectory prediction model are determined based on the converter framework; The encoder module for obtaining the initial variational adaptive transformer is based on the integration of the preset multi-head self-attention module and the preset adaptive graph convolutional network module. The encoder module acquires the time of the preset training sample sequence and the correlation information of each channel in the encoder, so as to output the dynamic hidden state of the training sample sequence based on the correlation information. Based on the dynamic hidden state, the decoder module of the initial variational adaptive transformer is reconstructed; wherein, the decoder module is a probabilistic generative model; The encoder module and the decoder module are combined based on the positions of the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer.
9. A non-volatile storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions include: The target robot acquires image data of the target to be inspected using one or more pre-installed acquisition devices. The location coordinates of the target to be inspected are obtained based on the image data, and the trajectory data sequence of the target to be inspected is determined according to the acquisition time corresponding to the location coordinates of the target to be inspected. The trajectory data sequence is input into a trajectory prediction model trained based on a preset variational adaptive transformer for prediction, so as to obtain the trajectory prediction result of the inspection target; wherein, the preset variational adaptive transformer is constructed based on a transformer framework, and the preset variational adaptive transformer includes: a multi-head self-attention module, an adaptive graph convolutional network module, and a probability generation module. Before inputting the trajectory data sequence into the trajectory prediction model trained based on a preset variational adaptive transformer for prediction, the following steps are also included: The components of the initial variational adaptive transformer for the trajectory prediction model are determined based on the converter framework; The encoder module for obtaining the initial variational adaptive transformer is based on the integration of the preset multi-head self-attention module and the preset adaptive graph convolutional network module. The encoder module acquires the time of the preset training sample sequence and the correlation information of each channel in the encoder, so as to output the dynamic hidden state of the training sample sequence based on the correlation information. Based on the dynamic hidden state, the decoder module of the initial variational adaptive transformer is reconstructed; wherein, the decoder module is a probabilistic generative model; The encoder module and the decoder module are combined based on the positions of the components of the initial variational adaptive transformer to obtain the preset variational adaptive transformer.
Citation Information
Patent Citations
Molecule generation and optimization method based on variational auto-encoder
CN114038516A
Word embedding with disentangling prior
US20220335216A1