A Pedestrian Action Prediction Method Based on a Dynamic Graph Convolution Attention Model

Through the dynamic graph convolution attention model, the pedestrian skeleton graph motion sequence is divided into historical and future sequences, and the features are extracted and correlation features are calculated, which solves the oversmooth problem in the deep graph convolution model, improving the effect of pedestrian motion prediction.

CN118522070BActive Publication Date: 2025-06-10SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410571490.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-06-10
Estimated Expiration
2044-05-07

AI Technical Summary

Technical Problem

Deep graph convolutional neural networks are prone to oversmoothing problems in pedestrian movement prediction, resulting in a decrease in prediction ability.

Method used

The pedestrian movement prediction method based on the dynamic graph convolution attention model is adopted. By obtaining the pedestrian skeleton graph motion sequence, it is divided into historical sequences and future sequences, the feature information is extracted using the preset convolutional neural network, and the correlation characteristics between the historical sequences and future sequences are determined through the attention neural network, and finally prediction is made through the dynamic graph convolutional network.

Benefits of technology

It effectively solves the problem of oversmoothing in graph convolution neural networks, and improves the accuracy and effectiveness of pedestrian movement prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118522070B_ABST
    Figure CN118522070B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of computer vision, and provides a pedestrian action prediction method based on a dynamic graph convolutional attention model. The method divides the pedestrian skeleton graph action sequence into a historical sequence and a future sequence, calculates the correlation features between the historical sequence and the future sequence through an attention mechanism, and then fuses the pedestrian skeleton graph action sequence and the correlation features through a dynamic graph convolutional network to obtain the predicted pedestrian action. Since deep graph convolution has the problem of over-smoothing, the present system alleviates the over-smoothing problem by introducing a dynamic graph convolutional network, making the pedestrian action prediction more accurate and capable of improving the effect of pedestrian action prediction ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a pedestrian action prediction method based on a dynamic graph convolutional attention model. Background Art

[0002] In the field of computer vision, pedestrian action prediction is an important task and has great significance for autonomous driving and traffic safety. In recent years, human action prediction based on skeleton graphs has attracted wide attention. By using a graph convolutional network (GCN), effective feature extraction can be performed on the skeleton graph of the human body, thereby achieving more accurate action prediction. GCN allows the model to consider the relationships between nodes on the graph when processing the skeleton graph, enabling the model to better capture the key information of human actions and thus improving the prediction effect. However, deep GCN models often suffer from over-smoothing problems during prediction, resulting in a decline in prediction ability. Summary of the Invention

[0003] An embodiment of the present invention provides a pedestrian action prediction method based on a dynamic graph convolutional attention model, aiming to solve the over-smoothing problem existing in the graph convolutional neural network and improve the prediction effect.

[0004] In a first aspect, an embodiment of the present invention provides a pedestrian action prediction method based on a dynamic graph convolutional attention model, including the following steps:

[0005] Obtain a pedestrian skeleton graph motion sequence, divide the pedestrian skeleton graph motion sequence into multiple motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence;

[0006] Extract the feature information of the motion subsequences, historical sequences, and future sequences according to a preset neural network based on convolutional operations, and determine the correlation features between the historical sequences and the future sequences through an attention neural network;

[0007] Predict the future action sequence of the pedestrian through a dynamic graph convolutional network according to the correlation features and the pedestrian skeleton graph motion sequence.

[0008] Combined with the first aspect, in the second implementation manner of the first aspect, divide the pedestrian skeleton graph motion sequence into multiple motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence. The specific division method is as follows:

[0009] For the pedestrian skeleton graph motion sequence X T =[x 1 , x 2 ,..., x T , divide it into (T - M - F + 1) motion subsequences in chronological order

[0010] Define the first M frames of each motion subsequence as the historical sequence X history , and the last M frames of the input pedestrian skeleton graph motion sequence as the future sequence X future .

[0011] Among them, X T represents the input pedestrian skeleton graph motion sequence, T is the number of frames of the incoming pedestrian skeleton graph sequence, K is the number of joint points of the pedestrian, and F and M are parameters for adjusting the number of motion subsequences. Specifically, the values of M and F depend on the number of frames T of the incoming pedestrian skeleton graph sequence.

[0012] Combined with the first aspect, in the third implementation manner of the first aspect, the preset neural network based on convolution operation includes three pedestrian skeleton graph motion sequence feature extractors. According to the preset neural network based on convolution operation, extract the feature information of the motion subsequence, historical sequence, and future sequence, and determine the correlation feature between the historical sequence and the future sequence through the attention neural network, including:

[0013] Use the discrete cosine transform to perform time-domain to frequency-domain conversion on the motion subsequence to obtain the converted motion subsequence;

[0014] Use the preset neural network based on convolution operation to obtain the feature information of the converted motion subsequence, historical sequence, and future sequence;

[0015] Use the attention neural network to calculate the correlation between the historical sequence and the future sequence according to the frequency-domain information of the motion subsequence, historical sequence, and future sequence features, and obtain the correlation feature.

[0016] Combined with the first aspect, in the fourth implementation manner of the first aspect, predict the future action sequence of the pedestrian through the dynamic graph convolutional network according to the correlation feature and the pedestrian skeleton graph motion sequence, including:

[0017] Use the dynamic graph convolutional neural network with graph convolutional layers, Tanh, and skip connections to fuse the correlation feature with the pedestrian skeleton graph motion sequence to obtain the pedestrian skeleton graph fusion information;

[0018] Perform inverse discrete cosine transform on the pedestrian skeleton graph fusion information to obtain the predicted future action sequence of the pedestrian.

[0019] Combined with the fourth implementation manner of the first aspect, in the fifth implementation manner of the first aspect, after performing inverse discrete cosine transform on the pedestrian skeleton graph fusion information to obtain the predicted future action sequence of the pedestrian, the method further includes: smoothing the real future action sequence of the pedestrian through the smoothing formula to obtain the smoothed real future action, and using the smoothed real future action for guided prediction. The smoothing formula is expressed as:

[0020]

[0021] Among them, is the smoothed action of the i-th frame, and T h is the starting frame number of the pedestrian's future action, and T f is the ending frame number of the pedestrian's future action, and x k represents the pedestrian's future action of the k-th frame in the pedestrian's future action sequence.

[0022] Combined with the fourth implementation manner of the first aspect, in the sixth implementation manner of the first aspect, the dynamic graph convolutional attention model includes a dynamic graph convolutional network, and the forward propagation formula of the dynamic graph convolutional network is expressed as:

[0023]

[0024]

[0025] Among them, H (l+1) represents the feature of the (l + 1)-th intermediate hidden layer, and α and β are coefficients for adjusting the distribution of feature information. represents the feature of the l-th layer of the (n + 1)-th residual block, which contains the features of the previous layer, and the proportion of historical feature information is adjusted by the coefficient γ.

[0026] In a second aspect, an embodiment of the present invention provides a pedestrian action prediction device based on a graph convolutional attention model, including:

[0027] An action division unit, configured to obtain a pedestrian skeleton graph motion sequence, divide the pedestrian skeleton graph motion sequence into multiple motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence;

[0028] An association calculation unit, configured to extract feature information of the pedestrian skeleton graph motion sequence, historical sequence, and future sequence according to a preset neural network based on convolutional operations, and determine the association features between the historical sequence and the future sequence through an attention neural network;

[0029] An action prediction unit, configured to predict the future action sequence of the pedestrian through a dynamic graph convolutional network according to the association features and the pedestrian skeleton graph motion sequence.

[0030] In a third aspect, an embodiment of the present invention provides a computer device, including: a processor, and a memory storing computer program instructions, and the processor reads and executes the computer program instructions to perform the steps of the pedestrian action prediction method based on graph convolution attention in the first aspect or any implementation manner of the first aspect.

[0031] Fourthly, an embodiment of the present invention provides a computer storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the steps of the pedestrian action prediction method based on graph convolutional attention in the first aspect or any implementation manner of the first aspect are realized.

[0032] In the pedestrian action prediction method based on graph convolutional attention according to the embodiment of the present invention, the method includes: dividing the pedestrian skeleton graph motion sequence into multiple motion subsequences, and each motion subsequence is divided into a historical sequence and a future sequence; extracting the features of the historical sequence and the future sequence through a preset neural network based on convolutional operations, and obtaining the correlation features between the historical sequence and the future sequence through an attention neural network; using the obtained correlation features, combining with the pedestrian skeleton graph motion sequence, and predicting the future action sequence of the pedestrian through a dynamic graph convolutional network, which can solve the over-smoothing problem existing in the graph convolutional neural network, and can improve the prediction ability effect by predicting the future action sequence of the pedestrian through the dynamic graph convolutional network. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0034] Figure 1 is a schematic flowchart of a pedestrian action prediction method based on a dynamic graph convolutional attention model provided by an embodiment of the present invention;

[0035] Figure 2 is a schematic diagram of the network structure of a specific example of a pedestrian action prediction method based on a dynamic graph convolutional attention model provided by an embodiment of the present invention;

[0036] Figure 3 is a schematic flowchart of the attention-based correlation feature extraction provided by an embodiment of the present invention;

[0037] Figure 4 is a structure diagram of a dynamic graph convolutional network provided by an embodiment of the present invention;

[0038] Figure 5 is a training flowchart of a dynamic graph convolutional network provided by an embodiment of the present invention;

[0039] Figure 6 is a pedestrian action prediction device based on a dynamic graph convolutional attention model provided by an embodiment of the present invention;

[0040] Figure 7 is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without some of these specific details. The following description of the embodiments is only provided to provide a better understanding of the present invention by showing examples of the present invention.

[0042] To solve the problems of the prior art, an embodiment of the present invention provides a pedestrian action prediction method based on a dynamic graph convolutional attention model. First, the pedestrian action prediction method based on the dynamic graph convolutional attention model provided by the embodiment of the present invention will be introduced below.

[0043] Figure 1 A flowchart showing a pedestrian action prediction method based on a dynamic graph convolutional attention model provided by an embodiment of the present invention is shown. As Figure 1 shown, the method may include the following steps:

[0044] S110. Obtain a pedestrian skeleton graph motion sequence, divide the pedestrian skeleton graph motion sequence into multiple motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence.

[0045] The pedestrian skeleton graph motion sequence is the motion sequence captured by a camera, and may include information such as the 3D skeleton coordinates of the pedestrian and the pedestrian motion time. The motion subsequence is a part of the input pedestrian skeleton graph motion sequence, which is a continuous sequence set, and the information contained is the same as that of the human skeleton graph motion sequence, but is shorter in time. The historical sequence is the sequence with an earlier time in the motion subsequence, which is a sequence set, and the future sequence is the last sequence of the pedestrian skeleton graph motion sequence.

[0046] The user captures the skeleton data of the pedestrian through a specific camera such as a Kinect, and inputs the skeleton data of the pedestrian through an input device at the terminal to obtain the input pedestrian skeleton graph motion sequence. Dividing the pedestrian skeleton graph motion sequence into a future sequence and a historical sequence according to the time sequence is beneficial to making full use of the historical information of the pedestrian skeleton graph motion sequence.

[0047] S120. Extract the feature information of the pedestrian skeleton graph motion sequence, the historical sequence, and the future sequence according to a preset neural network based on convolutional operations, and determine the correlation feature between the historical sequence and the future sequence through an attention neural network.

[0048] The feature information is the graph information obtained after convolution, which may include the speed information, joint position information, etc. of the pedestrian skeleton graph motion sequence. The correlation feature is the correlation score between the historical sequence and the future sequence calculated by the attention network. The preset neural network based on convolution operation extracts the feature information through convolution calculation, and the attention neural network calculates the correlation feature according to the feature information obtained after convolution according to a specific formula.

[0049] Specifically, the discrete cosine transform is used to perform time-domain to frequency-domain conversion on the pedestrian skeleton graph motion sequence, and the preset neural network based on convolution operation is used to obtain the feature information of the pedestrian skeleton graph motion sequence and the divided historical sequence and future sequence. Through the attention neural network, according to the obtained feature information, the correlation between the historical sequence and the future sequence is calculated to obtain the correlation feature.

[0050] S130. According to the correlation feature and the pedestrian skeleton graph motion sequence, predict the future action sequence of the pedestrian through the dynamic graph convolutional network.

[0051] Specifically, a dynamic graph convolutional neural network with a graph convolutional layer, Tanh, and skip connection wires is used to fuse the correlation feature and the pedestrian skeleton graph motion sequence, and the predicted future action sequence of the pedestrian is obtained through the inverse discrete cosine transform.

[0052] It can be seen that in the embodiment of the present application, by obtaining the input pedestrian skeleton graph motion sequence, dividing the pedestrian skeleton graph motion sequence into multiple motion subsequences, and further dividing each motion subsequence into a historical sequence and a future sequence, extracting the feature information of the pedestrian skeleton graph motion sequence, as well as the divided historical sequence and future sequence according to the preset neural network based on convolution operation, and obtaining the correlation feature between the historical sequence and the future sequence through the attention neural network, and predicting the future action sequence of the pedestrian through the dynamic graph convolutional network in combination with the pedestrian skeleton graph motion sequence according to the correlation feature, the prediction ability effect can be improved.

[0053] Figure 2 Shows the network structure of a specific example of a pedestrian action prediction method based on a dynamic graph convolutional attention model. The following is combined with Figures 2 to 5 Further elaborate on the embodiments of the present invention:

[0054] I. Action division process

[0055] The input pedestrian skeleton graph motion sequence is divided into subsequences in a certain order. The specific division method is as follows:

[0056] For the pedestrian skeleton graph motion sequence X T =[x 1 , x 2 ,..., x TDivide into (T - M - F + 1) motion subsequences in chronological order

[0057] Define the first M frames of each motion subsequence as the historical sequence X history , and the last M frames of the input pedestrian skeleton graph motion sequence as the future sequence X future .

[0058] Among them, X T represents the input pedestrian skeleton graph motion sequence, T is the number of frames of the incoming pedestrian skeleton graph sequence, K is the number of joint points of the pedestrian, and F and M are parameters for adjusting the number of motion subsequences. Specifically, the values of M and F depend on the number of frames T of the incoming pedestrian skeleton graph sequence.

[0059] II. Calculation process of the correlation characteristics between the historical sequence and the future sequence

[0060] Figure 3 Shows the process of the attention network calculating the correlation characteristics between the historical sequence and the future sequence. As Figure 3 shown, the method may include the following steps:

[0061] S310. The motion subsequence converts the time-domain information to the frequency domain through discrete cosine transform.

[0062] Among them, the discrete cosine transform is an information coding algorithm that can concentrate the expression of time-domain information, which is beneficial to subsequent feature extraction.

[0063] S320. Use a preset neural network based on convolutional operation to obtain the feature information of the converted motion subsequence, historical sequence, and future sequence.

[0064] Among them, the preset neural network based on convolutional operation includes three one-dimensional convolutional neural networks. The one-dimensional convolutional neural network is used to extract the motion sequence features, and the feature information of the motion sequence features can be obtained through convolutional operation.

[0065] S330. Calculate the correlation characteristics between the historical sequence and the future sequence through the attention neural network.

[0066] The formula for obtaining the feature information of the converted motion subsequence, historical sequence, and future sequence is expressed as:

[0067]

[0068] In the above formula, K, Q, and V are the feature information of the historical sequence, future sequence, and pedestrian skeleton graph motion subsequence respectively, and are calculated through a one-dimensional convolutional neural network. Among them, the pedestrian skeleton graph motion subsequence is first subjected to discrete cosine transform (DCT) to better represent the action information of the pedestrian.

[0069] The formula for calculating the correlation features of the historical sequence and the future sequence through the attention neural network is:

[0070]

[0071] att is the calculated correlation feature, T is the number of frames of the input pedestrian skeleton graph sequence, F and M are parameters for adjusting the number of motion subsequences when dividing actions, and K, Q, and V are the feature information of the historical sequence, the future sequence, and the pedestrian skeleton graph motion subsequence respectively.

[0072] III. Dynamic Graph Convolutional Network

[0073] Figure 4 The specific composition of the dynamic graph convolutional network is shown. The dynamic graph convolutional network includes at least one residual block, and the residual block is composed of multiple graph convolutional layers. There are skip connections inside the residual block, and there are also skip connections for the graph convolution between the residual blocks. The forward propagation formula of the dynamic graph convolutional network is expressed as:

[0074]

[0075]

[0076] Among them, H (l+1) represents the feature of the (l + 1)-th hidden layer in a residual block, and α and β are adjustments. represents the feature of the l-th layer of the (n + 1)-th residual block, which contains the features of the previous layer, and the proportion of the feature information is adjusted by the coefficient γ.

[0077] IV. Guided Prediction

[0078] Figure 5 The training process of the dynamic graph convolutional network is shown, including:

[0079] S510. Fuse the correlation features with the input pedestrian graph skeleton motion sequence through the graph convolutional layer;

[0080] S520. Multiple convolutional layers form a residual block, and there are skip connections inside and between the residual blocks, which can alleviate the over-smoothing problem;

[0081] S530. During training, after passing through each residual block, compare the intermediate prediction result with the smoothed true action;

[0082] S540. After passing through multiple residual blocks, finally output the predicted pedestrian action through the inverse discrete cosine transform.

[0083] After passing through each residual block, an intermediate prediction result will be output, which is compared with the smoothed true action to optimize the prediction performance of the model.

[0084] Among them, the adopted smoothing formula is:

[0085]

[0086] Among them, is the action after smoothing for the i-th frame, T h is the starting frame number of the action, T f is the ending frame number of the action, x k represents the action of the k-th frame.

[0087] The loss function is used to measure the error between the predicted value and the true value of the network model, so as to guide the learning of the network model parameters. The preset dynamic graph convolutional network is continuously learned based on the loss function, and the loss function formula of this neural network is expressed as:

[0088]

[0089]

[0090] Among them, represents the predicted action, J represents the true action, the subscript represents the i-th frame and the j-th joint point, T represents the time of the pedestrian motion sequence, and N represents the number of joints of the pedestrian skeleton graph. During the intermediate process represents the intermediate prediction result, J represents the smoothed true action, A represents the number of stacked residual blocks, L all represents the final loss function.

[0091] Figure 6 This is a pedestrian action prediction device 600 based on a dynamic graph convolutional attention model provided by an embodiment of the present invention, including:

[0092] An action division unit 610, configured to obtain a pedestrian skeleton graph motion sequence, divide the pedestrian skeleton graph motion sequence into multiple motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence;

[0093] An association calculation unit 620, configured to extract the feature information of the pedestrian skeleton graph motion sequence, historical sequence, and future sequence according to a preset neural network based on convolutional operations, and determine the association features between the historical sequence and the future sequence through an attention neural network;

[0094] An action prediction unit 630, configured to predict the future action sequence of the pedestrian through a dynamic graph convolutional network according to the association features and the pedestrian skeleton graph motion sequence.

[0095] As an implementable manner, the action division unit 610 is specifically configured to: for the pedestrian skeleton graph motion sequence X T =[x 1 , x2 ,..., x T is divided into (T - M - F + 1) motion subsequences in chronological order

[0096] Define the first M frames of each motion subsequence as the historical sequence X history , and the last M frames of the input pedestrian skeleton graph motion sequence as the future sequence X future .

[0097] Among them, X T represents the input pedestrian skeleton graph motion sequence, T is the number of frames of the incoming pedestrian skeleton graph sequence, K is the number of joint points of the pedestrian, and F and M are parameters for adjusting the number of motion subsequences.

[0098] As an implementable way, the association calculation unit 620 is specifically used for: performing time-domain to frequency-domain conversion on the motion subsequence by using discrete cosine transform to obtain the converted motion subsequence;

[0099] using a preset neural network based on convolutional operation to obtain the feature information of the converted motion subsequence, historical sequence, and future sequence;

[0100] using an attention neural network to calculate the correlation between the historical sequence and the future sequence according to the frequency-domain information of the motion subsequence, the feature information of the historical sequence, and the future sequence, and obtaining the correlation feature.

[0101] As an implementable way, the action prediction unit 630 is specifically used for: using a dynamic graph convolutional neural network with a graph convolutional layer, Tanh, and skip connections to fuse the correlation feature with the pedestrian skeleton graph motion sequence to obtain the pedestrian skeleton graph fusion information;

[0102] performing inverse discrete cosine transform on the pedestrian skeleton graph fusion information to obtain the predicted pedestrian future action sequence.

[0103] As an implementable way, after performing inverse discrete cosine transform on the pedestrian skeleton graph fusion information to obtain the predicted pedestrian future action sequence, the action prediction unit 630 is also specifically used for: guiding the prediction of the smoothed real future action through a smoothing formula, and the smoothing formula is expressed as:

[0104]

[0105] Among them, is the action after smoothing for the i-th frame, T h is the starting frame number of the action, T f is the ending frame number of the action, x k represents the action of the k-th frame.

[0106] As an implementable approach, the dynamic graph convolutional attention model includes a dynamic graph convolutional network, and the forward propagation formula of the dynamic graph convolutional network is expressed as:

[0107]

[0108]

[0109] Among them, H (l+1) represents the features of the (l + 1)-th intermediate hidden layer, and α and β are coefficients for adjusting the distribution of feature information. represents the features of the l-th layer of the (n + 1)-th residual block, which contains the features of the previous layer, and the proportion of historical feature information is adjusted by the coefficient γ.

[0110] Figure 7 FIG. shows the schematic hardware structure of the computer device provided by the embodiments of the present invention.

[0111] The computer device may include a processor 701 and a memory 702 storing computer program instructions.

[0112] Specifically, the above-mentioned processor 701 may include a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0113] The memory 702 may include a mass storage for data or instructions. By way of example and not limitation, the memory 702 may include a Hard Disk Drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. In one example, the memory 702 may include removable or non-removable (or fixed) media, or the memory 702 is a non-volatile solid-state memory. The memory 702 may be inside or outside the integrated gateway disaster recovery device.

[0114] In one example, the memory 702 may be a Read Only Memory (ROM). In one example, the ROM may be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), an Electrically Rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0115] The processor 701 reads and executes the computer program instructions stored in the memory 702 to implement Figure 1 the methods / steps S110 to S130 in the illustrated embodiment and achieve Figure 1 the corresponding technical effects achieved by the methods / steps executed by the illustrated example. For the sake of concise description, they will not be elaborated here.

[0116] In one example, the computer device may further include a communication interface 703 and a bus 710. Among them, as Figure 7 shown, the processor 701, the memory 702, and the communication interface 703 are connected through the bus 710 to complete the communication with each other.

[0117] The communication interface 703 is mainly used to implement the communication between the modules, devices, units, and / or devices in the embodiments of the present invention.

[0118] The bus 710 includes hardware, software, or both, and couples the components of the online data flow charging device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 710 may include one or more buses. Although the embodiments of the present invention describe and illustrate a specific bus, the present invention contemplates any suitable bus or interconnect.

[0119] This computer device can be used to execute the pedestrian action prediction method based on the dynamic graph convolutional attention model in the embodiments of the present invention, so as to implement the combination of Figure 1 and Figure 6 the pedestrian action prediction method and device based on the dynamic graph convolutional attention model described.

[0120] In addition, in combination with the method for predicting pedestrian actions based on the dynamic graph convolutional attention model in the above embodiments, an embodiment of the present invention can provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the methods for predicting pedestrian actions based on the dynamic graph convolutional attention model in the above embodiments is implemented.

[0121] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0122] It should also be noted that the functional blocks shown in the above structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0123] It also needs to be explained that the exemplary embodiments mentioned in the present invention describe some methods or devices based on a series of steps or devices. However, the present invention is not limited to the order of the above steps. That is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0124] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A pedestrian motion prediction method based on a dynamic graph convolutional attention model, characterized in that: The method comprises: Acquire a pedestrian skeleton image motion sequence, divide the pedestrian skeleton image motion sequence into a plurality of motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence; Extracting feature information of the motion subsequence, the historical sequence, and the future sequence according to a preset convolution operation-based neural network, and determining a correlation feature between the historical sequence and the future sequence through an attention neural network; According to the correlation feature and the motion sequence of the pedestrian skeleton image, the future action sequence of the pedestrian is predicted by a dynamic graph convolutional network, specifically including: using a dynamic graph convolutional neural network with a graph convolution layer, Tanh, and skip connection to fuse the correlation feature and the motion sequence of the pedestrian skeleton image to obtain pedestrian skeleton image fusion information; performing inverse discrete cosine transform on the pedestrian skeleton image fusion information to obtain a predicted pedestrian future action sequence; smoothing the real pedestrian future action sequence by a smoothing formula to obtain the smoothed real future action, and using the smoothed real future action to guide prediction, the smoothing formula is expressed as: in, is the smoothed motion of the i-th frame, T h is the starting frame number of the pedestrian’s future action, T f is the end frame number of the pedestrian's future action, x k represents the pedestrian's future action in the kth frame of the pedestrian's future action sequence; The dynamic graph convolution attention model includes a dynamic graph convolution network, and the forward propagation formula of the dynamic graph convolution network is expressed as: Among them, H (l+1) represents the features of the middle l+1 hidden layer, α and β are coefficients for adjusting the distribution of feature information, represents the features of the lth layer of the n+1th residual block, which contains the features of the previous layer, and the coefficient γ adjusts the proportion of historical feature information; The dynamic graph convolutional network is continuously learned based on the loss function, and the formula of the loss function is expressed as: in, represents the predicted action, J represents the real action, the subscript represents the i-th frame, the j-th joint point, T represents the time of the pedestrian motion sequence, N represents the number of joints in the pedestrian skeleton graph, and in the intermediate process represents the intermediate prediction result, J represents the smoothed real action, A represents the number of stacked residual blocks, and L all Represents the final loss function.

2. The pedestrian motion prediction method based on dynamic graph convolutional attention model according to claim 1, characterized in that: The step of dividing the pedestrian skeleton image motion sequence into a plurality of motion subsequences, and dividing each motion subsequence into a historical sequence and a future sequence, comprises: For the pedestrian skeleton motion sequence X T =[x 1 ,x 2 ,…,x T ] is divided into (TM-F+1) motion subsequences according to the time sequence Define the first M frames of each motion subsequence as the historical sequence X history , the last M frames of the input pedestrian skeleton motion sequence are the future sequence X future ; Among them, X T It represents the input pedestrian skeleton motion sequence, T is the number of frames of the input pedestrian skeleton sequence, K is the number of joints of the pedestrian, and F and M are parameters for adjusting the number of motion subsequences.

3. The pedestrian motion prediction method based on dynamic graph convolutional attention model according to claim 1, characterized in that: The extracting feature information of the motion subsequence, the historical sequence and the future sequence according to a preset convolution operation-based neural network, and determining the correlation feature between the historical sequence and the future sequence through an attention neural network, includes: Using discrete cosine transform to transform the motion subsequence from time domain to frequency domain to obtain a transformed motion subsequence; Acquire feature information of the converted motion subsequence, the historical sequence, and the future sequence using the preset convolution operation-based neural network; By using an attention neural network, the correlation between the historical sequence and the future sequence is calculated according to the frequency domain information of the motion subsequence, the feature information of the historical sequence and the future sequence, and the correlation feature is obtained.

4. A pedestrian motion prediction device based on a dynamic graph convolutional attention model, characterized in that: The device comprises: An action division unit is used to obtain a motion sequence of a pedestrian skeleton image, divide the motion sequence of the pedestrian skeleton image into a plurality of motion subsequences, and divide each motion subsequence into a historical sequence and a future sequence; A correlation calculation unit, configured to extract feature information of the pedestrian skeleton motion sequence, the historical sequence and the future sequence according to a preset convolution operation-based neural network, and determine a correlation feature between the historical sequence and the future sequence through an attention neural network; The action prediction unit is used to predict the future action sequence of the pedestrian through a dynamic graph convolution network according to the correlation feature and the motion sequence of the pedestrian skeleton image, specifically comprising: using a dynamic graph convolution neural network with a graph convolution layer, Tanh, and skip connection lines to fuse the correlation feature with the motion sequence of the pedestrian skeleton image to obtain pedestrian skeleton image fusion information; performing an inverse discrete cosine transform on the pedestrian skeleton image fusion information to obtain a predicted future action sequence of the pedestrian; smoothing the real future action sequence of the pedestrian through a smoothing formula to obtain the smoothed real future action, and using the smoothed real future action to guide the prediction, and the smoothing formula is expressed as: in, is the smoothed motion of the i-th frame, T h is the starting frame number of the pedestrian’s future action, T f is the end frame number of the pedestrian's future action, x k represents the pedestrian's future action in the kth frame of the pedestrian's future action sequence; The dynamic graph convolution attention model includes a dynamic graph convolution network, and the forward propagation formula of the dynamic graph convolution network is expressed as: Among them, H (l+1) represents the features of the middle l+1 hidden layer, α and β are coefficients for adjusting the distribution of feature information, represents the features of the lth layer of the n+1th residual block, which contains the features of the previous layer, and the coefficient γ adjusts the proportion of historical feature information; The dynamic graph convolutional network is continuously learned based on the loss function, and the formula of the loss function is expressed as: in, represents the predicted action, J represents the real action, the subscript represents the i-th frame, the j-th joint point, T represents the time of the pedestrian motion sequence, N represents the number of joints in the pedestrian skeleton graph, and in the intermediate process represents the intermediate prediction result, J represents the smoothed real action, A represents the number of stacked residual blocks, and L all Represents the final loss function.

5. A computer device, characterized in that: The device comprises: a processor, and a memory storing computer program instructions; The processor reads and executes the computer program instructions to implement a pedestrian motion prediction method based on a dynamic graph convolutional attention model as described in any one of claims 1-3.

6. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, which, when executed by a processor, implement a pedestrian motion prediction method based on a dynamic graph convolutional attention model as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Real-time action recognition method and system for multi-person scene

    CN112906545A

  • Video three-dimensional human body posture estimation method and system based on multistage supervision graph convolution

    CN114694261A