Trajectory prediction method and device

By introducing a double-headed attention mechanism into the graph attention network, the attention relationship between moving targets is solved, and the problem of difficulty in accurately predicting the motion trajectory in the multi-objective motion scenario in the prior art is solved, and the accuracy of trajectory prediction is improved.

CN114998388BActive Publication Date: 2025-05-16BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD

Patent Information

Application Number
CN202210657220.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-05-16
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

Existing trajectory prediction methods are difficult to accurately model the attention relationship between moving targets in multi-objective motion scenarios, affecting the prediction accuracy of the motion trajectory.

Method used

A double-headed attention mechanism is used to predict trajectory in the graph attention network, and the attention relationship between the moving targets is decomposed, and the attention of each node at the end point and during the movement is calculated separately.

Benefits of technology

The accuracy of trajectory prediction results is improved, and the characteristics and attention relationship of the moving target can be expressed more accurately, and the dynamic changes in motion scenarios can be adapted to dynamically changing motion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998388B_ABST
    Figure CN114998388B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a trajectory prediction method and device. A specific implementation of the method includes: determining a target relationship graph, the nodes in the target relationship graph represent moving targets, and the edges in the target relationship graph represent the attention relationship between moving targets; obtaining the historical motion trajectory of each node in the target relationship graph; using a pre-trained trajectory prediction model, based on a two-headed attention mechanism, according to the target relationship graph and the historical motion trajectory, to obtain the predicted motion trajectory of each node, the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the terminal position and a second attention calculation for each node's attention to other nodes at each moment during the motion process. This implementation helps to improve the accuracy of trajectory prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a trajectory prediction method and device. Background Art

[0002] Trajectory prediction has a wide range of applications in fields such as autonomous driving. With the development and application of deep learning technology, trajectory prediction methods based on recurrent neural networks and latent variable models have emerged. In order to model the attention of different moving targets to other moving targets, a method of trajectory prediction using a graph attention network based on an attention mechanism has emerged. Current trajectory prediction methods usually model the spatial correlation in the historical trajectory sequence of moving targets, but in actual multi-target motion scenarios, the attention relationship between moving targets will affect the motion trajectory of the moving targets. Summary of the invention

[0003] The embodiments of the present disclosure provide a trajectory prediction method and device.

[0004] In a first aspect, an embodiment of the present disclosure provides a trajectory prediction method, the method comprising: determining a target relationship graph, wherein nodes in the target relationship graph represent moving targets, and edges in the target relationship graph represent attention relationships between moving targets; obtaining the historical motion trajectory of each node in the target relationship graph; using a pre-trained trajectory prediction model, based on a two-headed attention mechanism, obtaining a predicted motion trajectory of each node according to the target relationship graph and the historical motion trajectory, wherein the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the end position and a second attention calculation for each node's attention to other nodes at each moment during the motion process.

[0005] In a second aspect, an embodiment of the present disclosure provides a trajectory prediction device, which includes: a determination unit, configured to determine a target relationship graph, wherein the nodes in the target relationship graph represent moving targets, and the edges in the target relationship graph represent attention relationships between moving targets; an acquisition unit, configured to acquire the historical motion trajectory of each node in the target relationship graph; a prediction unit, configured to use a pre-trained trajectory prediction model to obtain a predicted motion trajectory of each node based on the target relationship graph and the historical motion trajectory based on a two-headed attention mechanism, wherein the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the end position and a second attention calculation for each node's attention to other nodes at each moment during the motion process.

[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0008] The trajectory prediction method and device provided by the embodiments of the present disclosure perform trajectory prediction through a graph attention network, and divide the attention between moving nodes into the attention of each node to other nodes at the end position and the attention of each node to other nodes at each moment during the movement process, so that attention calculation can be performed more accurately in combination with different angles, thereby helping to improve the accuracy of trajectory prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0010] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0011] Figure 2 is a flow chart of an embodiment of a trajectory prediction method according to the present disclosure;

[0012] Figure 3 is a schematic diagram of an embodiment of a node partitioning method;

[0013] Figure 4 is a schematic diagram of an embodiment of graph attention computation;

[0014] Figure 5 is a schematic structural diagram of an embodiment of a trajectory prediction device according to the present disclosure;

[0015] Figure 6 It is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0017] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0018] Figure 1 An exemplary architecture 100 is shown to which an embodiment of a trajectory prediction method or a trajectory prediction apparatus of the present disclosure can be applied.

[0019] like Figure 1 As shown, the system architecture 100 may include moving targets in various scenes (such as vehicles 101, pedestrians 102, etc.) and a server 103. The arrows in the figure may indicate the direction of movement of the moving target. The dotted line in the figure may indicate the historical movement trajectory of the moving target. The server may consider the various moving targets in the scene in a graph, the nodes in the graph may represent the moving targets, and the edges in the graph may represent various relationships between the nodes, and then the server may predict the movement trajectory of each moving target in the scene based on methods such as the graph attention network, and then the server may also assist some moving targets (such as unmanned vehicles, etc.) in driving decisions based on the prediction results of the movement trajectory of each moving target.

[0020] It should be noted that the server 103 can be hardware or software. When the server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server 103 is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0021] It should be noted that the trajectory prediction method provided in the embodiments of the present disclosure is generally executed by the server 103 , and accordingly, the trajectory prediction device is generally disposed in the server 103 .

[0022] It should also be noted that the system architecture 100 may also include various terminal devices, and the server 103 may be connected to the terminal device in a communication manner such as a wired or wireless communication link or an optical fiber cable. The server 103 may send the obtained prediction result of the motion trajectory of each moving target to the terminal device.

[0023] The terminal device can be hardware or software. When the terminal device is hardware, it can be various electronic devices, including but not limited to smart phones, tablet computers, vehicle-mounted terminals, laptop computers, desktop computers, etc. When the terminal device is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0024] Various client applications can be installed on the terminal device, such as web browser applications, shopping applications, search applications, instant messaging tools, social platform software, navigation applications, map applications, etc.

[0025] It should also be pointed out that trajectory prediction applications may also be installed in the terminal device, and the terminal device may also predict the motion trajectory of each moving target based on the trajectory prediction application. At this time, the trajectory prediction method may also be executed by the terminal device, and accordingly, the trajectory prediction device may also be set in the terminal device.

[0026] It should be understood that Figure 1 The number of moving targets and servers in the example is only for illustration. Any number of moving targets and servers may be provided as required.

[0027] Continue to refer Figure 2 , which shows a process 200 of an embodiment of a trajectory prediction method according to the present disclosure. The trajectory prediction method comprises the following steps:

[0028] Step 201, determine the target relationship graph.

[0029] In this step, the nodes in the target relationship graph may represent moving targets. Moving targets may be various objects that can move. For example, moving targets include but are not limited to pedestrians, vehicles, robots, etc. There may be different moving targets in different scenarios. For example, in the scenario of road traffic analysis, moving targets may include pedestrians and vehicles. In some cases, vehicles may include but are not limited to: manned vehicles and unmanned vehicles, etc.

[0030] The edges in the target relationship graph can represent the attention relationship between the moving targets. The attention (or degree of attention) of a moving target "A" to another moving target "B" can represent the influence of the moving target "B" on the moving trajectory of the moving target "A". Specifically, the attention relationship can be set according to the actual application scenario. For example, the attention relationship can indicate whether there is attention between two moving targets. If there is no edge between the two moving targets, that is, there is no attention relationship, it can be said that the two moving targets do not have a significant impact on each other's moving trajectories. For another example, the attention relationship can indicate whether there is attention between the two moving targets, and the attention weight between the two moving targets with an attention relationship. Among them, the attention weight can indicate the degree of influence of one moving target on the moving trajectory of another moving target.

[0031] It should be noted that in some scenarios, moving targets may include moving targets of multiple categories. The attention relationship between two moving targets of different categories may be asymmetric. As an example, for two vehicles "V1" and "V2", since both belong to the same vehicle category, if "V1" has attention to "V2", then "V2" usually has the same attention to "V1". For a vehicle "V" and a pedestrian "P", since the two belong to different categories, the conditions for generating their attention may be different. Therefore, there may be no correlation between whether "V" has attention to "P" and whether "P" has attention to "V". At this time, the target relationship graph can be a directed graph.

[0032] In this embodiment, the execution subject of the trajectory prediction method (such as Figure 1 The server 103 shown in the figure can use various methods to determine the target relationship graph. Since the position of each moving target at each moment is constantly changing, and different positional relationships may have different attention, the target relationship graph constructed at each moment can be the same or different. The execution subject can also determine the target relationship graph at the required moment according to the needs.

[0033] Generally speaking, for any two nodes in the target relationship graph, it is possible to determine whether there is an attention relationship, that is, whether there is an edge, between the two nodes based on the positional relationship between the moving targets represented by the two nodes and the attention conditions set in advance for each node, thereby constructing the target relationship graph.

[0034] Among them, the positional relationship can be flexibly expressed in various ways. For example, the positional relationship can be expressed using the distance between the moving targets represented by two nodes (such as Euclidean distance, cosine distance, etc.). The position of the moving target represented by each node can also be flexibly expressed in various ways. For example, the position of the moving target can be represented by various methods such as coordinates, longitude and latitude. Specifically, a target relationship graph at any time can be constructed according to demand. At this time, the attention condition can be less than a preset distance threshold. Generally, at the prediction time, a target relationship graph can be constructed based on the current position of each node.

[0035] It should be noted that the construction of the target relationship graph can be implemented by the above-mentioned execution subject, and can also be implemented by other electronic devices. In this case, the execution subject can obtain the constructed target relationship graph from other electronic devices.

[0036] Step 202: Obtain the historical movement trajectory of each node in the target relationship graph.

[0037] In this embodiment, the historical motion trajectory of each node may be composed of the positions of the corresponding moving target at each time before the prediction time. For example, the historical motion trajectory of each node may be composed of the coordinates of the corresponding moving target at each time before the prediction time arranged in order from early to late. The above-mentioned execution subject may obtain the historical motion trajectory of each node from a local or other storage device (such as a vehicle-mounted terminal, a smart camera on the road, etc.).

[0038] Step 203, using the pre-trained trajectory prediction model, based on the dual-headed attention mechanism, according to the target relationship graph and the historical motion trajectory, obtain the predicted motion trajectory of each node.

[0039] In this embodiment, the trajectory prediction model can be constructed based on the Graph Attention Network (GAT). The Graph Attention Network mainly aggregates the neighbor nodes of each node through the Attention Mechanism to achieve adaptive allocation of attention weights for different neighbors, which can improve the expression ability of each node feature, thereby improving the accuracy of the trajectory prediction results based on each node feature. Among them, the Graph Attention Network and the Attention Mechanism are currently widely studied and applied known technologies, which will not be repeated here.

[0040] In this embodiment, the dual-headed attention mechanism may include a first attention calculation for each node's attention to other nodes at the end position and a second attention calculation for each node's attention to other nodes at each moment during the movement process. For each node, other nodes may refer to other nodes in the target relationship graph except the node, but since each node has an attention relationship with only some of the surrounding nodes in many cases, other nodes may refer to neighbor nodes of the node in the target relationship graph, and neighbor nodes refer to nodes that have an attention relationship with the node.

[0041] The first attention calculation may refer to the attention calculation performed at the moment when the node reaches the terminal position according to the attention relationship between the nodes at that moment. The second attention calculation may refer to the attention calculation performed at each moment during the node movement according to the attention relationship between the nodes at that moment. The specific attention calculation method can adopt the calculation method provided by the existing attention mechanism, which will not be repeated here.

[0042] As an example, feature aggregation of neighbor nodes in attention calculation can be implemented based on the following formula:

[0043] Enc v =LSTM(X)

[0044]

[0045] Enc e =Atten(Enc v , E e )

[0046] Enc x =Concat(Enc v ,Enc e )

[0047] Among them, "X" represents the vector of the historical motion trajectory of each neighbor node. For example, the historical motion trajectory of each neighbor node can be summed into a vector by using summation-pooling or other methods. v " represents the feature vector of the historical motion trajectory of the neighboring nodes extracted using LSTM (Long Short-Term Memory). t " represents the adjacency matrix of the target relationship graph at time "t". "K add ” and “K remove " are two preset convolution kernels, which aggregate the adjacency matrices at different times through the convolution of the sliding window in time, thereby helping to enhance the temporal information representation in the target relationship graph. "*" represents convolution. "E i " represents the category of the neighbor node "i" (such as pedestrians, vehicles, etc.), and different preset values ​​can be used to represent different categories. "·" represents the product. "E e " represents the attention relationship characteristics between nodes in the target relationship graph. "Atten(Enc v , E e )" means "Enc v " as "Query", "E e "As the attention calculation of "Key", "Enc e " is the attention calculation result. "Concat()" represents the concatenation operation. "Enc x ” represents the feature aggregation result of neighbor nodes.

[0048] Since the first attention calculation is for the node at the end position, and the second attention calculation is for the node at each moment in the motion process, the dual-headed attention mechanism implemented by the first attention calculation and the second attention calculation can not only perform attention analysis between nodes from the dimension of the entire process of the motion trajectory of each node, but also perform attention analysis between nodes from the dimension of the motion trajectory of each node at each moment, which helps to improve the accuracy of the output results of the attention mechanism, thereby improving the ability to express the characteristics of the moving target represented by the node.

[0049] The trajectory prediction model can more accurately express the characteristics of each node in the target relationship graph based on the dual-headed attention mechanism, and then predict the future movement trajectory of each node based on the characteristics of each node and the historical movement trajectory of the node, thereby obtaining the predicted movement trajectory of each node. The specific prediction method can adopt various existing trajectory prediction methods, which will not be repeated here.

[0050] The trajectory prediction model can be trained in advance by technicians using training samples based on machine learning methods. The training samples can be obtained from various actual sports scenes or from third-party data platforms. The training method of neural network based on machine learning is a well-known technology that is currently widely studied and applied, and will not be described in detail here.

[0051] Since the adoption of the above-mentioned dual-headed attention mechanism including the first attention calculation and the second attention calculation can improve the ability to express the characteristics of the moving target, the trajectory prediction model can also more accurately predict the motion trajectory of the moving target based on the better feature expression of the moving target, which helps to improve the accuracy of the trajectory prediction results.

[0052] In some optional implementations of this embodiment, the first attention calculation can be implemented by the following steps:

[0053] Step 1: Predict the final position of each node based on the historical motion trajectory.

[0054] In this step, various existing prediction methods can be used to predict the terminal position of each node. For example, for each node, the movement direction and average movement speed of the node are determined according to the historical movement trajectory of the node, and then the terminal position of the node is determined based on the movement direction and average movement speed.

[0055] Optionally, a Gaussian mixture model (GMM) can be used to predict the end position of each node based on the historical movement trajectory of the node. The Gaussian mixture model uses a Gaussian probability density function (normal distribution curve) to accurately quantify things, so as to decompose one thing into several models based on the Gaussian probability density function.

[0056] By describing the uncertainty of the node's position change by probability, the node's trajectory can be predicted more accurately. Therefore, the Gaussian mixture model can be used to predict the distribution of the endpoint coordinates of each node to more accurately represent the endpoint coordinates of each node.

[0057] Specifically, various existing prediction methods based on Gaussian mixture models can be used to predict the terminal position of each node according to the needs. As an example, assuming that the terminal coordinate distribution of the node is a two-dimensional Gaussian distribution, a multi-layer perceptron (MLP) can be used to build a fully connected network to predict the terminal coordinates of the node, and a Gaussian mixture model can be used to represent the terminal coordinates of the node. For details, please refer to the following formula. The input of the fully connected network can be the feature vector of the extracted historical motion trajectory of the node, and the output is the distribution parameter of the terminal coordinates of the node:

[0058]

[0059] Among them, “I” is a vector representing the historical motion trajectory of the node to be predicted. “X” is a vector representing the historical motion trajectory of other nodes (such as neighboring nodes of the node to be predicted, etc.). x " represents the feature vector representing the node to be predicted after attention calculation based on the target relationship graph at the prediction moment. "des" represents the distribution parameter of the endpoint coordinates of the node. "p" represents the probability of the endpoint coordinates. "Z" is a hidden variable.

[0060] Step 2: Determine a first matrix based on the positional relationship between the end positions of each node.

[0061] In this step, each row in the first matrix may correspond to each node in the target relationship graph, and each column may also correspond to each node in the target relationship graph. The elements in the first matrix may represent the attention weights between the nodes. It should be noted that, since the attention conditions between different types of moving targets may be different, the attention weights of two nodes to each other may be the same or different.

[0062] Specifically, various methods can be used to determine the first matrix. For example, for each moving target indicated by a node, a corresponding relationship between the positional relationship between the moving target and other moving targets and the attention weight can be set. At this time, the corresponding attention weight can be queried according to the positional relationship between the nodes to obtain the first matrix.

[0063] Step three, use the first matrix to perform the first attention calculation.

[0064] In this step, the attention weights between the nodes in the first matrix can be used to perform aggregation operations on the neighboring nodes of each node to improve the expression ability of each node feature. The specific calculation method can be flexibly set according to application requirements.

[0065] By calculating the attention through the distance between the predicted end points of each moving target, it is possible to represent the attention relationship that always exists between the moving targets during the movement process to a certain extent, that is, to achieve a long-term attention modeling.

[0066] In some optional implementations of this embodiment, the second attention calculation may include the following steps:

[0067] Step 1: At each moment in the process of node movement, determine the positional relationship between the nodes.

[0068] In this step, the positional relationship between each node at each moment can be determined. Among them, the positional relationship between the nodes can be determined based on the position of each node at each moment. The position of each node at each moment can be determined by various methods according to the specific application scenario. For example, each node can collect its own position at each moment and upload it to the execution subject. For another example, the execution subject can monitor the movement process of each node and collect the position of each node at each moment.

[0069] Step 2: Determine the second matrix based on the position relationship between the nodes at each moment.

[0070] In this step, each row in the second matrix may correspond to each node in the target relationship graph, and each column may also correspond to each node in the target relationship graph. The elements in the second matrix may represent the attention weights between the nodes. It should be noted that, since the attention conditions between different types of moving targets may be different, the attention weights of two nodes to each other may be the same or different. In addition, since the positions of the nodes are different at different times, the attention weights between the nodes may be the same or different at different times.

[0071] Specifically, various methods can be used to determine the second matrix. For example, for the moving target indicated by each node, the corresponding relationship between the positional relationship between the moving target and other moving targets and the attention weight can be set. At this time, the corresponding attention weight can be queried according to the positional relationship between each node to obtain the second matrix.

[0072] Step three, use the second matrix to perform the second attention calculation.

[0073] In this step, at each moment, the attention weights between the nodes in the second matrix can be used to perform aggregation operations on the neighboring nodes of each node to improve the expression ability of each node feature at each moment. The specific calculation method can be flexibly set according to application requirements.

[0074] By calculating the attention of each moving target to other nodes at each moment in the movement process, the dependence of each moving target on each other at each moment can be expressed to a certain extent, that is, a short-term attention modeling is realized.

[0075] In some optional implementations of this embodiment, the elements in the first matrix and the second matrix may be determined by the following steps:

[0076] It should be noted that the following steps can be used to determine the elements in the first matrix, and can also be used to determine the elements in the second matrix.

[0077] Step 1: For an element representing the attention weight of a first node to a second node, obtain an attention distance threshold of the first node to the second node.

[0078] In this step, the first node and the second node may be nodes represented by the row and column where the element to be determined is located, respectively. The attention threshold may be used to determine whether there is an attention relationship between nodes. The attention distance threshold may be pre-set by a technician according to an actual application scenario, or determined according to historical experimental data. The attention distance threshold of the first node to the second node may be used to determine whether the first node has attention to the second node.

[0079] Step 2: determine a calculation result with the natural constant as the base and the negative number of the distance between the first node and the second node as the exponent.

[0080] In this step, the distance between the first node and the second node may be various distances, such as Euclidean distance, cosine distance, and the like.

[0081] Step three, in response to determining that the operation result is greater than the attention distance threshold, determining the operation result as the element value.

[0082] In this step, if the calculation result is greater than the attention distance threshold of the first node to the second node, the obtained calculation result can be used as the value of the element representing the attention weight of the first node to the second node.

[0083] Step 4: In response to determining that the operation result is not greater than the attention distance threshold, determining that the value of the element is zero.

[0084] In this step, if the calculation result is not greater than the attention distance threshold of the first node to the second node, "0" can be used as the value of the element representing the attention weight of the first node to the second node.

[0085] As an example, the following formula may be used to determine the value of an element representing the attention weight of the first node to the second node:

[0086] f(x, THRESHOLD) = e -x , ife -x >THRESHOLD

[0087] =0, otherwise

[0088] Among them, "x" represents the distance between nodes. "THRESHOLD" represents the attention distance threshold. "f(x, THRESHOLD)" represents the element value. "e" is a natural constant.

[0089] Generally speaking, the closer the distance between moving targets, the greater the attention weight, and moving targets usually have a certain degree of selectivity. For example, according to expected behavior, they pay attention to the surrounding moving targets related to themselves, while ignoring some other moving targets, that is, they no longer pay attention to some moving targets. Therefore, we can set an attention threshold to screen moving targets with attention relationships, and use an exponential function to characterize the relationship between attention and distance between moving targets, so as to achieve a more accurate quantification of attention.

[0090] In some optional implementations of the present embodiment, the positional relationship between the terminal positions of each node may include the Euclidean distance between the terminal positions of each node. The positional relationship between each node at each moment may include the Euclidean distance and / or cosine distance between the positions of each node at each moment. For example, the Euclidean distance between the terminal positions of each node may be used in the above-mentioned first attention calculation process, and in the above-mentioned second attention calculation process, the Euclidean distance and cosine distance between the positions of each node at each moment may be used to perform attention calculations respectively.

[0091] Optionally, the second matrix may be determined by the following steps:

[0092] Step 1: Determine the first sub-matrix based on the Euclidean distance between nodes at each moment.

[0093] In this step, the Euclidean distance between nodes can be used as the positional relationship between the nodes, and the first sub-matrix can be determined using the above method for determining the second matrix.

[0094] Step 2: Determine the second sub-matrix based on the cosine distance between the nodes at each moment.

[0095] In this step, the cosine distance between the nodes may be used as the positional relationship between the nodes, and the second sub-matrix may be determined using the above method for determining the second matrix.

[0096] Step 3: Determine the second matrix using the first sub-matrix and the second sub-matrix.

[0097] In this step, after obtaining the first submatrix and the second submatrix, various fusion methods can be used to fuse the first submatrix and the second submatrix to obtain the second matrix. For example, the fusion method includes but is not limited to: weighted summation, Hadamard Product, etc.

[0098] In practical applications, the Euclidean distance or cosine distance between nodes can be used to perform attention calculation according to the application scenario or actual needs, which can improve the flexibility of the calculation process.

[0099] In some optional implementations of this embodiment, the attention calculation included in the above-mentioned dual-headed attention mechanism also includes: for each node, other nodes can be divided into three node sets according to the positional relationship between other nodes and the node. Among them, the three node sets can correspond to the front, side and back of the node respectively. Then, attention calculations can be performed on the three node sets respectively to obtain the attention calculation results corresponding to the three node sets respectively, and then the attention calculation results corresponding to the three node sets respectively are combined as the attention calculation results of the node.

[0100] Specifically, the nodes located in front of the node may form a first node set, the nodes located on the side of the node may form a second node set, and the nodes located on the back of the node may form a third node set. The specific ranges corresponding to the front, side, and back of each node may be flexibly determined in advance according to the actual application scenario, and then the node set to which each other node belongs may be determined by judging the range of the position of each other node.

[0101] For example, at each moment, the velocity vector of each target can be calculated based on the historical motion trajectory of each node, and then the cosine distance between any two nodes can be calculated using the velocity vector, thereby dividing the other nodes near each node into three node sets. Figure 3 , which shows a schematic diagram of an embodiment of a node partitioning method. Figure 3 As shown, for node 301, the nodes within 120 degrees on the front side with its movement direction as the center line can be organized into the first type of nodes (such as node 302 in the figure), the nodes within 60 degrees on the rear side can be organized into the second type of nodes (such as nodes 303 and 304 in the figure), and the nodes within the side range except the front and rear sides can be organized into the third type of nodes (such as node 305 in the figure).

[0102] For each node, after dividing other nodes into three types of nodes, attention calculations can be performed for each type of node separately, and then the attention calculation results corresponding to each type can be aggregated through various methods such as multi-layer perceptrons to obtain the attention calculation results of the node. In the aggregation process, different weights can be set for different types.

[0103] Generally, different types of moving targets may have different attention to each other. Moreover, for each moving target, the moving target generally pays different attention to moving targets in different ranges around it. For example, a moving target usually pays relatively more attention to the moving target in front of it, and relatively less attention to the moving target on its side, and relatively less attention to the moving target behind it. Based on this, for each moving target, by classifying the moving targets around it, such as dividing the neighboring nodes based on the perspective of the moving target, to perform attention calculations on different types of moving targets separately, and then aggregating the attention calculation results, it helps to further improve the flexibility and accuracy of attention calculations.

[0104] In some optional implementations of this embodiment, the predicted motion trajectory of each node may include at least two predicted motion trajectories. The above trajectory prediction model may also be used to output the probability corresponding to each predicted motion trajectory of each node.

[0105] Through multiple distributed motion trajectory predictions, it is more conducive to expressing the uncertainty of trajectory prediction and facilitating the presentation of motion trajectory evolution under different circumstances.

[0106] In some optional implementations of the present embodiment, the attention calculation included in the above-mentioned dual-headed attention mechanism may also include: first using the self-attention mechanism to perform self-attention calculation to obtain a self-attention calculation result, and then fusing the calculation result obtained by the first attention calculation with the self-attention calculation result to obtain a first fusion result, and fusing the calculation result obtained by the second attention calculation with the self-attention calculation result to obtain a second fusion result, and then aggregating the first fusion result and the second fusion result as the output result of the dual-headed attention mechanism.

[0107] As an example, the attention calculation process included in the above two-headed attention mechanism is shown as follows:

[0108] A des (Y des )=f(Distance(Y des ))

[0109] A dis (Y δ-1 )=f(Distance(Y δ-1 ))

[0110] A angle (Y δ-1 )=f(Angle(Y δ-1 ))

[0111]

[0112]

[0113]

[0114] Y δ-1 =MLP(A1·(Y δ-1 W V ), A2·(Y δ-1 W′ V ))

[0115] f(x, THRESHOLD) = e -x , if e -x >THRESHOLD

[0116] =0, otherwise

[0117] Among them, "x" represents the distance between nodes. "THRESHOLD" represents the attention distance threshold. "f(x, THRESHOLD)" represents the element value. "e" is a natural constant. "Y des ” and “Y δ-1 "A" represents the matrix composed of the positions of each node at the end time and the moment before the predicted time. des ”, “A dis ” and “A angle " respectively represent the first matrix calculated at the end time and the first sub-matrix and the second sub-matrix calculated at the moment before the prediction time. "W Q ”,"W K ”,"W V ”,"W V `" and "A attention " respectively represent the Q parameter matrix, K parameter matrix, different V parameter matrices and the attention weight matrix obtained by attention calculation. "A 1 attention ” and “A 2 attention " respectively represent the self-attention calculation results at the moment before the prediction moment and the end moment. "A 1 ” and “A 2 " represents the attention calculation results at the moment before the prediction moment and the end moment respectively. "Softmax" represents the activation function. "Cat" represents the connection operation. "() T” denotes a transpose operation. “⊙” denotes a Hadamard product. “MLP” denotes a multi-layer perceptron.

[0118] Continue to see Figure 4 ,in, Figure 4 A schematic diagram of an embodiment of graph attention calculation is shown. Figure 4 As shown in the figure, i " represents the feature vector of the node. The solid line represents feature aggregation, and multiple solid lines between nodes correspond to multiple heads in the multi-head attention calculation. The dotted line indicates that no feature aggregation is performed. "concat" and "avg" are the fusion calculation methods such as merging or averaging of examples. Figure 4 As an example, the self-attention calculation result of “h1” and “h i The attention calculation results of "h2", "h4" and "h5" are respectively fused to obtain the attention calculation result of h1, that is, the new feature representation of "h1" is obtained as "h / 1". The "α ij " represents the attention weight of node "j" to node "i". "exp" and "Softmax j " are respectively used as examples to calculate "α ij "The exponential function and activation function used. "h i ” and “h j ” represent the feature vectors of node “i” and node “j” respectively. “W Q ” and “W K " are the weights of node "i" and node "j" respectively. When "α ij When ” is 0, it means removing the “j” node, that is, the features of the “j” node are not aggregated when performing attention calculation on the “i” node, such as nodes “h3” and “h6” in the figure.

[0119] Since the attention of moving targets has many characteristics, such as paying attention to multiple different types of moving targets at the same time, different attention to different moving targets, different attention allocation in different scenarios, remembering the movement status of some moving targets even if attention is lost, selecting surrounding moving targets to allocate attention according to expected behaviors, unknown factors that affect attention changes, etc., in order to model the attention of moving targets, the present application scheme divides the attention of moving targets into long-term attention based on the predicted terminal position corresponding to the moving targets and short-term attention based on the real-time distance between moving targets, and combines it with the self-attention mechanism. It can more accurately aggregate the attention of each node (i.e., each moving target) to its neighboring nodes, thereby more accurately expressing the characteristics of each node, and can adapt to the dynamically changing motion scenes of neighboring nodes, thereby more accurately and flexibly predicting the motion trajectory of each moving target.

[0120] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a trajectory prediction device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0121] like Figure 5 As shown, the trajectory prediction device 500 provided in this embodiment includes a determination unit 501, an acquisition unit 502 and a prediction unit 503. The determination unit 501 is configured to determine a target relationship graph, wherein the nodes in the target relationship graph represent moving targets, and the edges in the target relationship graph represent the attention relationship between moving targets; the acquisition unit 502 is configured to obtain the historical motion trajectory of each node in the target relationship graph; the prediction unit 503 is configured to use a pre-trained trajectory prediction model, based on a dual-headed attention mechanism, to obtain the predicted motion trajectory of each node according to the target relationship graph and the historical motion trajectory, wherein the trajectory prediction model is constructed based on a graph attention network, and the dual-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the terminal position and a second attention calculation for each node's attention to other nodes at each moment during the motion process.

[0122] In this embodiment, in the trajectory prediction device 500, the specific processing of the determination unit 501, the acquisition unit 502 and the prediction unit 503 and the technical effects thereof can be referred to in Figure 2 The relevant descriptions of step 201, step 202 and step 203 in the corresponding embodiment are not repeated here.

[0123] In some optional implementations of the present embodiment, the above-mentioned first attention calculation includes the following steps: predicting the end position of each node based on the historical motion trajectory; determining a first matrix based on the positional relationship between the end positions of each node, wherein the elements in the first matrix represent the attention weights between the nodes; and performing the first attention calculation using the first matrix.

[0124] In some optional implementations of this embodiment, the above-mentioned first attention calculation includes the following steps: according to the historical motion trajectory, using a Gaussian mixture model to predict the end position of each node.

[0125] In some optional implementations of this embodiment, the above-mentioned second attention calculation includes the following steps: determining the positional relationship between each node at each moment in the node movement process; determining a second matrix based on the positional relationship between each node at each moment, wherein the elements in the second matrix represent the attention weights between the nodes; and performing the second attention calculation using the second matrix.

[0126] In some optional implementations of this embodiment, the positional relationship includes the distance between nodes; and the values ​​of the elements in the first matrix and the second matrix are determined by the following steps: for an element representing the attention weight of the first node to the second node, obtain an attention distance threshold of the first node to the second node, wherein the first node and the second node are nodes represented by the row and column where the element is located; determine an operation result with a natural constant as the base and the negative of the distance between the first node and the second node as the exponent; in response to determining that the operation result is greater than the attention distance threshold, determine the operation result as the element value; in response to determining that the operation result is not greater than the attention distance threshold, determine that the element value is zero.

[0127] In some optional implementations of this embodiment, the positional relationship between the endpoint positions of each node includes the Euclidean distance between the endpoint positions of each node, and the positional relationship between each node at each moment includes the Euclidean distance and / or cosine distance between the positions of each node at each moment.

[0128] In some optional implementations of the present embodiment, the positional relationship between the nodes at each moment includes the Euclidean distance and the cosine distance between the positions of the nodes at each moment; and the second matrix is ​​determined based on the positional relationship between the nodes at each moment, including: determining the first sub-matrix based on the Euclidean distance between the nodes at each moment; determining the second sub-matrix based on the cosine distance between the nodes at each moment; and determining the second matrix using the first sub-matrix and the second sub-matrix.

[0129] In some optional implementations of the present embodiment, the attention calculation included in the above-mentioned dual-headed attention mechanism also includes: for each node, dividing the other nodes into three node sets according to the positional relationship between the other nodes and the node, wherein the three node sets correspond to the front, side and rear of the node respectively; performing attention calculation on the three node sets respectively to obtain the attention calculation results corresponding to the three node sets respectively, and combining the attention calculation results corresponding to the three node sets respectively as the attention calculation result of the node.

[0130] In some optional implementations of this embodiment, the predicted motion trajectory of each node includes at least two predicted motion trajectories; and the trajectory prediction model is further used to output the probability corresponding to each predicted motion trajectory of each node.

[0131] In some optional implementations of the present embodiment, the attention calculation included in the above-mentioned dual-headed attention mechanism also includes: using the self-attention mechanism to perform self-attention calculation to obtain a self-attention calculation result; fusing the calculation result obtained by the first attention calculation with the self-attention calculation result to obtain a first fusion result; fusing the calculation result obtained by the second attention calculation with the self-attention calculation result to obtain a second fusion result; aggregating the first fusion result and the second fusion result as the output result of the dual-headed attention mechanism.

[0132] The device provided by the above-mentioned embodiment of the present disclosure determines a target relationship graph through a determination unit, wherein the nodes in the target relationship graph represent moving targets, and the edges in the target relationship graph represent the attention relationship between the moving targets; the acquisition unit acquires the historical motion trajectory of each node in the target relationship graph; the prediction unit uses a pre-trained trajectory prediction model to obtain the predicted motion trajectory of each node based on the target relationship graph and the historical motion trajectory based on a two-headed attention mechanism, wherein the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the terminal position and a second attention calculation for each node's attention to other nodes at each moment during the movement process, thereby performing trajectory prediction through the graph attention network, and dividing the attention between the moving nodes into the attention of each node at the terminal position to other nodes and the attention of each node to other nodes at each moment during the movement process, so as to combine different angles for more accurate attention calculation, thereby helping to improve the accuracy of the trajectory prediction results.

[0133] Reference below Figure 6 , which shows an electronic device (eg, Figure 1 The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The terminal device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0134] like Figure 6As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0135] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0136] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0137] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0138] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: determines a target relationship graph, wherein nodes in the target relationship graph represent moving targets, and edges in the target relationship graph represent attention relationships between moving targets; obtains the historical motion trajectory of each node in the target relationship graph; uses a pre-trained trajectory prediction model, based on a two-headed attention mechanism, to obtain the predicted motion trajectory of each node according to the target relationship graph and the historical motion trajectory, wherein the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the end position and a second attention calculation for each node's attention to other nodes at each moment during the motion process.

[0139] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0140] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0141] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units may also be arranged in a processor, for example, may be described as: a processor includes a determination unit, an acquisition unit, and a prediction unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases, for example, the acquisition unit may also be described as "a unit for acquiring the historical motion trajectory of each node in the target relationship graph".

[0142] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A trajectory prediction method, comprising: Determine a target relationship graph, wherein nodes in the target relationship graph represent moving targets, and edges in the target relationship graph represent attention relationships between moving targets; Obtain the historical movement trajectory of each node in the target relationship graph; Using a pre-trained trajectory prediction model, a predicted motion trajectory of each node is obtained according to the target relationship graph and the historical motion trajectory based on a two-headed attention mechanism, wherein the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the terminal position and a second attention calculation for each node's attention to other nodes at each moment during the motion process.

2. The method according to claim 1, wherein: The first attention calculation comprises the following steps: According to the historical motion trajectory, predict the terminal position of each node; Determine a first matrix based on the positional relationship between the end positions of each node, wherein the elements in the first matrix represent the attention weights between the nodes; A first attention calculation is performed using the first matrix.

3. The method according to claim 2, wherein: The step of predicting the terminal position of each node according to the historical motion trajectory includes: According to the historical motion trajectory, a Gaussian mixture model is used to predict the terminal position of each node.

4. The method according to claim 2, wherein: The second attention calculation comprises the following steps: At each moment during the node movement, determine the positional relationship between the nodes; Determine a second matrix based on the positional relationship between the nodes at each moment, wherein the elements in the second matrix represent the attention weights between the nodes; A second attention calculation is performed using the second matrix.

5. The method according to claim 4, wherein: The positional relationship includes the distance between nodes; and The values ​​of the elements in the first matrix and the second matrix are determined by the following steps: For an element representing an attention weight of a first node to a second node, obtaining an attention distance threshold of the first node to the second node, wherein the first node and the second node are nodes represented by the row and column where the element is located; Determine a result of an operation with a natural constant as a base and a negative number of a distance between the first node and the second node as an exponent; In response to determining that the operation result is greater than the attention distance threshold, determining the operation result as an element value; In response to determining that the operation result is not greater than the attention distance threshold, determining that the element takes a value of zero.

6. The method according to claim 5, wherein: The positional relationship between the end positions of the nodes includes the Euclidean distance between the end positions of the nodes, and the positional relationship between the nodes at each moment includes the Euclidean distance and / or cosine distance between the positions of the nodes at each moment.

7. The method according to claim 6, wherein: The positional relationship between the nodes at each moment includes the Euclidean distance and the cosine distance between the positions of the nodes at each moment; as well as The determining of the second matrix based on the positional relationship between the nodes at each moment includes: Determine a first submatrix based on the Euclidean distance between each node at each moment; Determine a second submatrix based on the cosine distance between the nodes at each moment; A second matrix is ​​determined using the first sub-matrix and the second sub-matrix.

8. The method according to claim 7, wherein: The attention calculation included in the dual-head attention mechanism also includes: For each node, according to the positional relationship between the other nodes and the node, the other nodes are divided into three node sets, wherein the three node sets correspond to the front, side and rear of the node respectively; Attention calculations are performed on the three node sets respectively to obtain attention calculation results corresponding to the three node sets respectively, and the attention calculation results corresponding to the three node sets respectively are combined as the attention calculation result of the node.

9. The method according to claim 3, wherein: The predicted motion trajectory of each node includes at least two predicted motion trajectories; and The trajectory prediction model is also used to output the probability corresponding to each predicted motion trajectory of each node.

10. The method according to claim 1, wherein: The attention calculation included in the dual-head attention mechanism also includes: Use the self-attention mechanism to perform self-attention calculation to obtain the self-attention calculation result; Fusing the calculation result obtained by the first attention calculation with the self-attention calculation result to obtain a first fusion result; Fusing the calculation result obtained by the second attention calculation with the self-attention calculation result to obtain a second fusion result; Aggregate the first fusion result and the second fusion result as the output result of the dual-head attention mechanism.

11. A trajectory prediction device, comprising: A determination unit is configured to determine a target relationship graph, wherein nodes in the target relationship graph represent moving targets, and edges in the target relationship graph represent attention relationships between moving targets; An acquisition unit, configured to acquire a historical motion trajectory of each node in the target relationship graph; The prediction unit is configured to use a pre-trained trajectory prediction model to obtain a predicted motion trajectory of each node based on the target relationship graph and the historical motion trajectory based on a two-headed attention mechanism, wherein the trajectory prediction model is constructed based on a graph attention network, and the two-headed attention mechanism includes a first attention calculation for each node's attention to other nodes at the terminal position and a second attention calculation for each node's attention to other nodes at each moment during the motion process.

12. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.

13. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Method for intelligently selecting sensing service node in vehicle networking

    CN108600938A

  • Vehicle trajectory prediction method based on environmental attention neural network model

    CN112215337A

Cited By

  • Target trajectory prediction method and system with domain generalization ability

    CN116912661A

  • A target trajectory prediction method and system with domain generalization capability

    CN116912661B