Method, device and electronic device for determining motion intention

By using object detection and heat map technology in intelligent driving, the problem of reduced target detection accuracy in intelligent driving is solved, and the accurate judgment of the movement intention of obstacles is achieved, and the misjudgment rate is reduced.

CN118366119BActive Publication Date: 2025-05-16CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410361161.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-05-16
Estimated Expiration
2044-03-27

AI Technical Summary

Technical Problem

In intelligent driving, the accuracy of target detection decreases when the obstacle is close to the bicycle, making it difficult to accurately determine the driving intention of the obstacle, and misjudgment is prone to occur.

Method used

By acquiring the image to be analyzed, object detection is performed to obtain the predicted position box and category of the moving target, the initial heat map contains the key point information of the moving target, match the object detection results and the heat map, adjust the heat map size to obtain an adaptive heat map, and analyze the key points in the heat map to determine the motion intention.

Benefits of technology

The target detection accuracy when the obstacle is close to the bicycle is improved, and the movement intention of the moving target is accurately determined, reducing misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118366119B_ABST
    Figure CN118366119B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device and electronic device for determining motion intention, and relates to the field of image processing technology, and is intended to accurately determine the motion intention of a moving target. The method includes: acquiring an image to be analyzed; performing target detection on the image to be analyzed to obtain a target detection result; generating at least one initial thermal map based on the image to be analyzed; matching the target detection result with the initial thermal map to obtain a moving target corresponding to each initial thermal map; adjusting the size of the initial thermal map corresponding to each moving target according to the target detection result to obtain an adaptive thermal map; analyzing the key points of the moving target in the image to be analyzed according to the adaptive thermal map corresponding to each moving target to obtain the motion intention of each moving target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device and electronic device for determining motion intention. Background Art

[0002] During intelligent driving, obstacles are identified through target detection, and then the key points of the obstacles are detected to determine the obstacle's driving intention. However, when the obstacle is close to the vehicle, the accuracy of target detection decreases, and it is difficult to accurately determine the key points of the obstacle. Therefore, it is easy to misjudge the obstacle's driving intention. Summary of the invention

[0003] In order to overcome the problems existing in the related art, the present disclosure provides a method, device and electronic device for determining motion intention. The technical solution of the present disclosure is as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a method for determining a motion intention is provided, comprising:

[0005] Acquire an image to be analyzed; the image to be analyzed includes at least one moving target;

[0006] Performing target detection on the image to be analyzed to obtain a target detection result; the target detection result includes: a predicted position frame of the moving target and a predicted category of the moving target;

[0007] Generate at least one initial thermal map according to the image to be analyzed; the initial thermal map includes key point information of the moving target;

[0008] Matching the target detection result with the initial heat map to obtain a moving target corresponding to each initial heat map;

[0009] According to the target detection result, the size of the initial heat map corresponding to each of the moving targets is adjusted to obtain an adaptive heat map; the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map;

[0010] According to the adaptive heat map corresponding to each of the moving targets, the key points of the moving targets in the image to be analyzed are analyzed to obtain the movement intention of each of the moving targets.

[0011] Optionally, analyzing key points of the moving targets in the image to be analyzed according to the adaptive heat map corresponding to each moving target to obtain the moving intention of each moving target includes:

[0012] According to the adaptive heat map corresponding to each of the moving targets, extracting a local image corresponding to each of the adaptive heat maps from the image to be analyzed, and adjusting the size of the local image to the size of the adaptive heat map corresponding to the local image;

[0013] According to the adaptive heat map corresponding to the local image, weighted processing is performed on the local image to obtain a weighted image;

[0014] According to the key points of the moving target in the weighted image, the direction angle of the moving target is calculated to obtain the moving direction intention of the moving target.

[0015] Optionally, adjusting the size of the initial heat map corresponding to each of the moving targets according to the target detection result to obtain an adaptive heat map includes:

[0016] According to the predicted category of the moving target, obtaining a correlation between the size of the position box of each moving target of the predicted category and the distance of the moving target;

[0017] Determining the distance of the moving target according to the association relationship and the size of the predicted position box of the moving target;

[0018] According to the distance of the moving target, the size of the initial heat map is adjusted to obtain the adaptive heat map.

[0019] Optionally, generating at least one initial thermal map according to the image to be analyzed includes:

[0020] Acquire a feature map of the image to be analyzed;

[0021] Performing sliding convolution on the feature map to obtain a response value of each point of the feature map;

[0022] According to the response values ​​of each point of the feature map, the feature map is divided into at least one feature sub-map, and the feature sub-map containing the key point is determined as a local feature map;

[0023] Determine the point with the largest response value in the local feature map as the key point of the moving object;

[0024] Generate a response map according to the response value of the local feature map, and activate each of the response maps respectively to obtain a heat map to be filtered;

[0025] According to the distance between each point in the local feature map and the key point in the local feature map, Gaussian filtering is performed on the thermogram to be filtered to obtain the initial thermogram.

[0026] Optionally, in the case where the image to be analyzed includes a plurality of continuous images, analyzing the key points of the moving target in the image to be analyzed according to the adaptive heat map corresponding to each moving target to obtain the movement intention of each moving target includes:

[0027] Tracking a moving target in each of the images to be analyzed;

[0028] For each of the moving targets, the motion attributes of the moving target are determined according to the adaptive thermal map corresponding to the moving target in each of the images to be analyzed and the key points of the moving target in the images to be analyzed; the motion attributes include: direction angle, speed, acceleration, historical motion trajectory and predicted motion trajectory;

[0029] The movement intention of the movement target is determined according to the movement attribute of the movement target.

[0030] Optionally, tracking a moving target in each of the images to be analyzed includes:

[0031] Establishing identifiers for key points of the moving target in a target image, and recording position information and feature information of the key points; the target image is: the first image among the multiple images to be analyzed;

[0032] In the event that the key point is lost in any of the images to be analyzed, predicting the position of the key point;

[0033] Matching and filtering the same key points in each of the images to be analyzed to obtain key point trajectories;

[0034] According to the key point trajectory, the position of the moving target in each of the images to be analyzed is determined.

[0035] Optionally, performing target detection on the image to be analyzed to obtain a target detection result includes:

[0036] Acquire a feature map of the image to be analyzed;

[0037] Inputting the feature map into a region proposal network to obtain at least one candidate region, and determining a probability that the candidate region is the moving target;

[0038] Predicting the candidate area to obtain a plurality of initial position frames, probabilities of the initial position frames containing the moving target, and predicted categories of the moving target in the initial position frames;

[0039] According to the probability that the moving target is contained in the initial position frame, non-maximum suppression is performed on the initial position frame whose overlap ratio exceeds an overlap threshold to obtain the predicted position frame.

[0040] Optionally, performing target detection on the image to be analyzed to obtain a target detection result includes:

[0041] Inputting the image to be analyzed into the target detection network of the image processing model to obtain the target detection result;

[0042] The step of generating at least one initial thermal map according to the image to be analyzed includes:

[0043] Inputting the image to be analyzed into the heat map generation network of the image processing model to obtain the initial heat map;

[0044] The training steps of the image processing model include:

[0045] Inputting the image samples to be analyzed into the initial image processing model, obtaining target detection result samples and initial heat map samples; the image samples to be analyzed include moving target samples, and the moving target samples carry: position frame information, category information, position information of key point samples, and category information of the key point samples; the target detection result samples include predicted position frame information and predicted category information of the moving target samples; the initial heat map samples include predicted position information and predicted category information of the key point samples;

[0046] Establishing a target detection loss function according to the position frame information and category information of the moving target sample and the target detection result sample;

[0047] Establishing a heat map loss function according to the position information and category information of the key point samples, and the predicted position information and predicted category information of the key point samples;

[0048] Based on the target detection loss function and the heat map loss function, the initial image processing model is trained to obtain the trained image processing model.

[0049] According to a second aspect of an embodiment of the present disclosure, a device for determining a motion intention is provided, comprising:

[0050] An acquisition module, used for acquiring an image to be analyzed; the image to be analyzed includes at least one moving target;

[0051] A detection module, used to perform target detection on the image to be analyzed to obtain a target detection result; the target detection result includes: a predicted position frame of the moving target and a predicted category of the moving target;

[0052] A generating module, configured to generate at least one initial thermal map according to the image to be analyzed; the initial thermal map includes key point information of the moving target;

[0053] A matching module, used to match the target detection result with the initial heat map to obtain a moving target corresponding to each of the initial heat maps;

[0054] An adjustment module, used to adjust the size of the initial heat map corresponding to each of the moving targets according to the target detection result to obtain an adaptive heat map; the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map;

[0055] The analysis module is used to analyze the key points of the moving targets in the image to be analyzed according to the adaptive heat map corresponding to each moving target, so as to obtain the movement intention of each moving target.

[0056] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method for determining motion intention as described in the first aspect.

[0057] According to a fourth aspect of an embodiment of the present disclosure, a non-volatile readable storage medium is provided. When instructions in the non-volatile readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method for determining motion intention as described in the first aspect.

[0058] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0059] In the disclosed embodiment, an initial heat map is generated based on the image to be analyzed, and the initial heat map includes key point information of the moving target, so the key points can be accurately determined. In addition, target detection is performed on the image to be analyzed to obtain a target detection result, and then the size of the initial heat map can be adjusted according to the target detection result to obtain an adaptive heat map of each moving target, wherein the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map. When analyzing the key points of the moving target in the image to be analyzed according to the adaptive heat map corresponding to the moving target, the weight of the moving target can be automatically adjusted, thereby effectively improving the accuracy of the determined moving target's motion intention.

[0060] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0062] Figure 1 is a flow chart of a method for determining a motion intention shown in an embodiment of the present disclosure;

[0063] Figure 2 is a schematic diagram of the architecture of the image processing model of the embodiment of the present disclosure;

[0064] Figure 3 is a schematic diagram of a flow chart of determining a motion intention in an embodiment of the present disclosure;

[0065] Figure 4 is a block diagram of a device for determining motion intention shown in an embodiment of the present disclosure;

[0066] Figure 5 Schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0067] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0068] The embodiment of the present disclosure provides a method for determining motion intention, which can be implemented by an electronic device, and the electronic device can be a terminal, a server, and other devices. Among them, the terminal can be an autopilot, a smart phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer (PC), and other devices; the server can be a single server or a server cluster composed of multiple servers. In some embodiments, the method can also be implemented by multiple electronic devices, for example, the method for determining the motion intention of the present disclosure can be implemented by multiple servers.

[0069] For example, the electronic device may be an autopilot and carried by the vehicle. The autopilot may communicate over the network and obtain photos of the vehicle's surroundings through sensors carried by the vehicle. The autopilot processes the photos of the vehicle's surroundings to determine moving targets around the vehicle and the movement intentions of each moving target.

[0070] A method for determining motion intention provided by an embodiment of the present disclosure can be applied to various technical fields, including but not limited to cloud technology, intelligent transportation, intelligent driving, maps, navigation and other technical fields.

[0071] Figure 1 is a flow chart of a method for determining a motion intention shown in an embodiment of the present disclosure, such as Figure 1 As shown, the method for determining the movement intention may include steps S11 to S16.

[0072] In step S11, an image to be analyzed is obtained.

[0073] The images to be analyzed in different application scenarios may be different. In the field of intelligent driving, the images to be analyzed may be images of the surroundings of the vehicle captured by a camera on the vehicle. The images to be analyzed include moving targets, which are targets that have the possibility of movement or the ability to move. For example, moving targets may be pedestrians, vehicles, and animals, but fixed pillars are not moving targets.

[0074] In step S12, target detection is performed on the image to be analyzed to obtain a target detection result.

[0075] The target detection can be performed on the image to be analyzed by a deep learning model suitable for the target detection task to obtain a target detection result. The target detection result includes: a predicted position frame of the moving target and a predicted category of the moving target.

[0076] Optionally, the deep learning model can be a model such as Faster R-CNN (Region-based Convolutional Neural Network), YOLO (You Only Look Once, a target detection model based on deep learning), SSD (Single Shot Multibox Detector, a single-stage target detector), etc. These models have a region proposal network (Region Proposal Network, RPN) and a classifier, which can simultaneously detect the location box of the moving target and the category of the moving target to obtain the target detection result.

[0077] Optionally, taking the target detection network as Faster R-CNN as an example, step S12 may include steps S121 to S124. The target detection network may include a convolutional network for image feature extraction, a region proposal network, a region interest pooling layer, and a target classification layer.

[0078] In step S121, a feature map of the image to be analyzed is obtained.

[0079] Convolutional Neural Networks (CNN) or Backbone Network can be used to extract features of the image to be analyzed to obtain a feature map of the image to be analyzed.

[0080] The image to be analyzed is input into the convolutional neural network, and features are extracted through the convolutional layer of the convolutional neural network. The convolutional layer uses filters to perform convolution operations on the image and extract different features. Pooling operations, such as maximum pooling or average pooling, are performed after the convolutional layer. The pooling layer can reduce the size of the feature map while retaining key information. An activation function is used to introduce nonlinearity and increase the expressive power of the network. The activation function can be a ReLU (Rectified Linear Unit) activation function. After multiple convolution and pooling operations, the output of the convolutional neural network is the feature map of the image to be analyzed. These feature maps contain abstract feature representations of the image to be analyzed at different levels, which can be used for subsequent classification, detection or other tasks.

[0081] In step S122, the feature map is input into a region proposal network to obtain at least one candidate region, and the probability that the candidate region is the moving target is determined.

[0082] The extracted feature map is input into the region proposal network, which generates multiple candidate regions for each window by sliding the window on the feature map, and predicts the probability that each candidate region is the target and the correction of the location box of the candidate region.

[0083] In step S123, the candidate region is predicted to obtain a plurality of initial position frames, probabilities of the initial position frames containing the moving target, and predicted categories of the moving targets in the initial position frames.

[0084] In order to make a more accurate prediction of the candidate regions, the candidate regions generated by the region proposal network are first sent to the region of interest pooling layer (RoIPool). The region interest pooling layer maps candidate regions of different sizes into feature maps of fixed size for subsequent processing.

[0085] The feature maps output by the region interest pooling layer are respectively input into the two fully connected layers of the target classification layer to obtain multiple initial position boxes, the probability of the initial position box containing the moving target, and the predicted category of the moving target in the initial position box. The two fully connected layers are used for target classification and position box regression respectively.

[0086] In step S124, according to the probability that the moving target is contained in the initial position frame, non-maximum suppression is performed on the initial position frame whose overlap ratio exceeds the overlap threshold to obtain the predicted position frame.

[0087] Multiple position frames may be predicted for one target. In order to obtain more accurate target detection results, non-maximum suppression can be performed on the initial position frames whose overlap ratio exceeds the overlap threshold. Only the initial position frames with the highest probability of containing the moving target among the initial position frames whose overlap ratio exceeds the overlap threshold are retained, and the retained initial position frames are determined as the predicted position frames of the moving target.

[0088] In the disclosed embodiment, Faster R-CNN introduces the RPN network, which can be trained end-to-end, and has high accuracy and efficiency in target detection. This method allows better use of deep learning technology to perform target detection tasks, and is more robust.

[0089] In step S13, at least one initial thermal map is generated according to the image to be analyzed.

[0090] The initial heat map includes key point information of the moving target.

[0091] The key points of the moving target are key points for judging the moving intention of the moving target. For example, when the moving target is a vehicle, the key points may be wheels; when the moving target is a pedestrian, the key points may be human feet.

[0092] The heatmap generation method is actually a feature matching technique by applying convolution kernels on feature maps.

[0093] Optionally, the initial heat map can be generated by the following method: obtaining a feature map of the image to be analyzed; performing sliding convolution on the feature map to obtain response values ​​of each point of the feature map; dividing the feature map into at least one feature sub-map according to the response values ​​of each point of the feature map, and determining the feature sub-map containing the key point as a local feature map; determining the point with the largest response value in the local feature map as the key point of the moving target; generating a response map according to the response value of the local feature map, and activating each of the response maps respectively to obtain a heat map to be filtered; performing Gaussian filtering on the heat map to be filtered according to the distance between each point in the local feature map and the key point in the local feature map to obtain the initial heat map.

[0094] The method for obtaining the feature map of the image to be analyzed can refer to the previous article. Use the convolution kernel to slide on the feature map, perform a convolution operation on each position of the feature map, and obtain the response value of the convolution kernel at that position. The feature map can be divided into multiple feature sub-maps according to the response values ​​of each point, and it can be determined which feature sub-maps are more likely to contain key points. The feature sub-map containing key points is the feature sub-map containing the moving target. The feature sub-map containing the key point is determined as a local feature map, and the point with the largest response value in the local feature map is determined as the key point of the moving target. The point with the largest response value in the local feature map can be determined by finding argmax (the value of the variable that makes a function reach the maximum value).

[0095] According to the response values ​​of each point in the local feature map, a response map can be generated, and each pixel point of the response map represents a response value. The response map is mapped to the range of 0 to 1 using an activation function to obtain an initial heat map containing key points. The activation function can be a sigmoid activation function.

[0096] The initial heat map can be adjusted by Gaussian filtering to make the distribution around the key point smoother. Assume that the key point μ = (μ x , μ y ), then the value of the point (x, y) on the heat map is:

[0097]

[0098] Among them, heatmap(x,y) represents the value of point (x,y) on the heat map, μ x is the horizontal coordinate of the key point, μ y is the ordinate of the key point, x is the abscissa of the point, y is the ordinate of the point, e is the natural logarithm, and σ represents the standard deviation of the Gaussian function.

[0099] It is difficult to accurately define a key point as a certain pixel position, so it is also complicated to accurately label key points. Points near key points often have similarities, and directly marking them as negative samples may introduce interference in network training. To alleviate this problem, using Gaussian function for "soft labeling" can make the training of neural networks smoother and help better convergence.

[0100] The introduction of Gaussian graphs can also provide directional guidance in network training. The closer to the target point, the greater the activation value. This design allows the network to quickly approach the key point in a more targeted manner. This clever use of Gaussian graphs enhances the network's ability to perceive small changes, helping to achieve more targeted and rapid convergence to the target point.

[0101] By adopting the technical solution of the embodiment of the present disclosure and moving the convolution kernel on the feature map plane, local information can be captured, local details can be focused on, and the positions of key points in the feature map can be accurately found. The generated initial heat map contains the key points of the moving target, which is helpful for judging the moving target's motion intention based on the key points. In addition, the generated initial heat map can be adjusted by the adaptive method described later to better adapt to the size and shape of the moving target.

[0102] Optionally, a neural network can be used to classify or locate the image content of the image to be analyzed, and the activation value (activation map) of the neural network can be visualized as an initial heat map. The image to be analyzed can be input into the selected convolutional neural network, and a specific layer of the neural network obtains a feature map or activation value. These feature maps contain abstract representations of the input image at different levels of the neural network, which can be understood as the degree of response of the network to different features. The activation value of the selected layer (or the activation value after specific processing) is processed, and the activation value is normalized, and then the processed activation value is visualized as an initial heat map. Among them, the point with the highest activation value after processing is determined as the key point information of the moving target.

[0103] In step S14, the target detection result is matched with the initial thermal map to obtain a moving target corresponding to each initial thermal map.

[0104] When generating the initial heat map, the category of the moving target in the initial heat map and the location of the moving target in the initial heat map can be preliminarily determined. In order to improve the accuracy of the subsequent judgment of the moving target's motion intention, the target detection result and the initial heat map can be matched to accurately determine the moving target corresponding to the initial heat map, and then determine the predicted category and predicted position box of the moving target in the initial heat map.

[0105] The target detection results and the initial heat map may be matched by calculating the similarity between the vector representation of the image area to be analyzed represented by the initial heat map and the vector representation of the image area corresponding to each target detection result.

[0106] In step S15, the size of the initial heat map corresponding to each of the moving targets is adjusted according to the target detection result to obtain an adaptive heat map.

[0107] The size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map.

[0108] According to the target detection result, the distance between the moving target corresponding to the target detection result and the camera can be determined. Specifically, according to the predicted category of the moving target, the association between the size of the position box of each moving target of the predicted category and the distance of the moving target is obtained; according to the association and the size of the predicted position box of the moving target, the distance of the moving target is determined.

[0109] The closer the moving target is, the larger the position frame is, and the farther the moving target is, the smaller the position frame is. Moving targets of the same distance but different categories may have different sizes, so the position frames of each category of moving targets need to be compared separately.

[0110] The association between the size of the position frame of different categories of moving targets and the distance of the moving targets can be determined in advance. After determining the predicted category of the moving target, the association corresponding to the predicted category is obtained. According to the position frame of the moving target, the association is queried to determine the distance of the moving target.

[0111] After determining the distance of the moving target, the size of the initial heat map can be adjusted according to the distance of the moving target to obtain the adaptive heat map. The size of the adaptive heat map is automatically adjusted according to the distance of the moving target included. The size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map.

[0112] On the basis of the above technical solution, the size of the initial heat map can be adjusted by adjusting the standard deviation of the Gaussian function, thereby obtaining an adaptive heat map. The size of the heat map is related to the standard deviation of the Gaussian function. By adjusting the value of the standard deviation of the Gaussian function, the size of the Gaussian circle can be changed, and moving targets of different sizes can be obtained at different distances, so that the size of the position box of the moving target is proportional to the standard deviation of the Gaussian function, so an adaptive heat map can be obtained.

[0113] Optionally, the initial heat map can also be resized through image processing or data visualization tools.

[0114] By adopting the technical solution of the embodiment of the present disclosure, the distance of the moving target can be determined through the target detection result of the moving target, and then the size of the initial heat map can be adaptively adjusted according to the distance of the moving target to obtain an adaptive heat map. In this way, based on the adaptive heat map obtained to analyze the image to be analyzed, the weight of the moving target can be automatically adjusted, effectively improving the accuracy of the determined moving target's movement intention.

[0115] In step S16, according to the adaptive heat map corresponding to each of the moving targets, the key points of the moving targets in the image to be analyzed are analyzed to obtain the movement intention of each of the moving targets.

[0116] The initial thermal map includes key point information of the moving target, so the adaptive thermal map corresponding to each moving target also includes key point information of the moving target. According to the key point information in the adaptive thermal map corresponding to each moving target, the key points of the moving target in the image to be analyzed can be determined, and then the key points of the moving target in the image to be analyzed can be analyzed to obtain the motion intention of each moving target.

[0117] Optionally, according to the adaptive heat map corresponding to each of the moving targets, the key points of the moving targets in the image to be analyzed are analyzed to obtain the movement intention of each of the moving targets, which may include: according to the adaptive heat map corresponding to each of the moving targets, extracting the local image corresponding to each of the adaptive heat maps from the image to be analyzed, and adjusting the size of the local image to the size of the adaptive heat map corresponding to the local image; according to the adaptive heat map corresponding to the local image, performing weighted processing on the local image to obtain a weighted image; according to the key points of the moving target in the weighted image, calculating the direction angle of the moving target to obtain the movement direction intention of the moving target.

[0118] According to the size of the adaptive heat map, the image area corresponding to the adaptive heat map in the image to be analyzed can be adjusted to the same size as the adaptive heat map. First, a local image corresponding to each adaptive heat map is extracted from the image to be analyzed, and then the size of the local image is adjusted to the size of the adaptive heat map corresponding to the local image.

[0119] The heat map can also characterize the importance of each image area of ​​the image to be analyzed. Therefore, the local image can be weighted according to the adaptive heat map to obtain a weighted image. Then, according to the key points of the moving target in the weighted image, the direction angle of the moving target is calculated to obtain the moving direction intention of the moving target point.

[0120] Taking the moving target as a vehicle as an example, the position information of the two key points where the vehicle wheels touch the ground can be obtained through adaptive heat maps. These key points may represent the contact points between the tires on one side of the vehicle and the ground, providing important clues about the position of the vehicle. Combined with the position information of the camera that took the image to be analyzed, the direction angle of the vehicle can be calculated. A common method is to use the principle of trigonometry. Specifically, the relative position vector between the two can be calculated through the position coordinates of the camera and the coordinates of the two points where the wheels on one side of the vehicle touch the ground. Then, using geometric calculation methods such as the inverse tangent function, the direction angle of the vehicle relative to the camera can be obtained.

[0121] This direction angle can provide information about the vehicle's current direction of travel. Such direction angle calculation has important applications in scenarios such as autonomous driving or traffic monitoring, helping to understand the vehicle's motion state, thereby better planning and adjusting the vehicle's driving strategy. By combining adaptive heat maps and geometric calculations, effective acquisition of vehicle direction information can be achieved.

[0122] In this way, by adjusting the size of the image area corresponding to the adaptive heat map in the image to be analyzed, each moving target can be adjusted to an appropriate size to avoid the image of the moving target being too large or too small, which affects the analysis of the motion intention, thereby improving the accuracy of the motion intention analysis. In addition, the direction angle of the moving target can also be determined through the information of the key points.

[0123] By adopting the technical solution of the embodiment of the present disclosure, an initial heat map is generated according to the image to be analyzed, and the initial heat map includes key point information of the moving target, so the key points can be accurately determined. In addition, target detection is performed on the image to be analyzed to obtain a target detection result, and then the size of the initial heat map can be adjusted according to the target detection result to obtain an adaptive heat map of each moving target, wherein the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map. When analyzing the key points of the moving target in the image to be analyzed according to the adaptive heat map corresponding to the moving target, the weight of the moving target can be automatically adjusted, thereby effectively improving the accuracy of the determined movement intention of the moving target.

[0124] Based on the above technical solution, the moving direction intention of the moving target can be obtained by analyzing only one image to be analyzed. When there are multiple consecutive images to be analyzed, the moving target's direction angle, speed, acceleration, historical movement trajectory, predicted movement trajectory and other movement attributes can be determined by target tracking, thereby determining the moving target's movement intention more accurately.

[0125] Specifically, in the case that the image to be analyzed includes multiple continuous images to be analyzed, target tracking can be performed on the moving target in each image to be analyzed; for each moving target, the motion attribute of the moving target is determined according to the adaptive thermal map corresponding to the moving target in each image to be analyzed, and the key points of the moving target in the image to be analyzed; the motion attributes include: direction angle, speed, acceleration, historical motion trajectory and predicted motion trajectory; according to the motion attributes of the moving target, the motion intention of the moving target is determined.

[0126] Tracking a moving target in each image to be analyzed may include: establishing an identifier for a key point of the moving target in the target image, and recording the position information and feature information of the key point; the target image is: the first image of multiple images to be analyzed; in the event that the key point is lost in any of the images to be analyzed, predicting the position of the key point; matching and filtering the same key point in each of the images to be analyzed to obtain a key point trajectory; and determining the position of the moving target in each of the images to be analyzed based on the key point trajectory.

[0127] In the first frame of the image to be analyzed, the key points of the moving target are detected, and then a unique identifier is established for each key point, and its position and feature information are recorded. This stage usually also includes building an initial motion model of the target.

[0128] The motion of the moving target can be estimated by the position change of the key points in consecutive frames. This can use methods such as optical flow, Kalman filter, particle filter, etc. to capture the dynamic changes of the moving target and update the current state of the moving target.

[0129] The key points can be re-detected in the new frame, and then the newly detected key points can be associated with the key points in the previous frame through a matching algorithm. Common association methods include nearest neighbor matching, optical flow matching, or feature descriptor-based matching to ensure that the key points of the moving target are correctly associated.

[0130] Filtering techniques (such as Kalman filtering) can be used to remove noise and instability to obtain smoother keypoint trajectories. This helps improve the stability of moving object tracking, especially in the presence of occlusion or dynamic environments.

[0131] The historical motion information of the moving target can be used to estimate the position of the key point in the next frame using a prediction model. This is important for dealing with temporary occlusion or disappearance of the moving target, and helps maintain the continuity of moving target tracking. It is understandable that in the case of the moving target being lost, the moving target may not be included in the image to be analyzed.

[0132] The image area where the moving target is located can be re-detected regularly to cope with possible detection failures, occlusions or scene changes. Re-detection helps to resume tracking of moving targets and ensure the robustness of the entire system in complex environments.

[0133] Through the above method, the moving target tracking system can accurately and stably track the key points obtained from the adaptive heat map in consecutive frames, improve the robustness of detection, and at the same time be able to handle the challenges of dynamic scenes and moving targets.

[0134] For each moving target, the motion attributes of the moving target can be determined based on the adaptive heat map corresponding to the moving target in the image to be analyzed and the key points of the moving target in the image to be analyzed; the motion attributes include: direction angle, speed, acceleration, historical motion trajectory and predicted motion trajectory; based on the motion attributes of the moving target, the motion intention of the moving target is determined.

[0135] The movement intention of the moving target can be judged by comprehensively considering the direction angle, speed, trajectory and other information of the key points.

[0136] The direction angle of the moving target can be calculated based on the key points. This can be achieved by calculating the movement direction of the key points in consecutive frames or the spatial distribution of the key points in one frame. The direction angle reflects the current direction of the dynamic moving target.

[0137] The motion information of key points can be used to estimate the speed of the moving target. By tracking the position changes of key points in consecutive frames, their relative speed can be calculated. Speed ​​information is crucial for understanding the behavior of moving targets.

[0138] Based on the velocity, the acceleration of the moving target can be calculated. Acceleration can be estimated by the rate of change of velocity, and acceleration is a key indicator for judging acceleration or deceleration behavior.

[0139] The trajectory of key points in consecutive frames can be analyzed, and behaviors such as lane changes and turns can be detected based on the trajectory of key point positions. For example, if a key point shows continuous left-right movement, it may indicate that the moving target is performing a turning operation.

[0140] Based on information such as direction angle, speed, acceleration, etc., the trajectory prediction model can be used to estimate the future trajectory of the moving target. This helps to better understand the expected actions of dynamic moving targets, including acceleration, deceleration, or steering.

[0141] Using the labeled training data, a machine learning classifier, such as a support vector machine (SVM) or a deep learning model, can be built to classify the driving intention of a moving target. A machine learning classifier can learn from multiple features and predict the specific behavior of a moving target.

[0142] Based on multiple images to be analyzed, the judgment of motion intention is continuously updated to adapt to possible changes in the moving target, which helps to capture the instantaneous behavior of the moving target more accurately.

[0143] By comprehensively considering multiple motion attributes such as direction, speed, acceleration, historical motion trajectory and predicted motion trajectory, the motion intention of the moving target can be judged more accurately, including different behaviors such as acceleration, deceleration, and steering, which helps the autonomous driving system interact with the moving target more safely.

[0144] Optionally, the method for determining motion intention of the embodiment of the present disclosure may be implemented based on an image processing model. Figure 2 Schematic diagram of the architecture of the image processing model of the embodiment of the present disclosure. The image processing model includes a target detection network and a heat map generation network. The target detection network can output the target detection result of the image to be analyzed, and the heat map generation network can output the initial heat map. The target detection network can include a convolutional network for image feature extraction, a region proposal network, a region interest pooling layer, and a target classification layer.

[0145] The image to be analyzed is input into the target detection network of the image processing model to obtain the target detection result; the image to be analyzed is input into the heat map generation network of the image processing model to obtain the initial heat map. The working principles of the target detection network and the heat map generation network can be referred to above.

[0146] After obtaining the target detection results and the initial heat map, the motion intention of the moving target can be determined by referring to the method described above, which will not be repeated here.

[0147] The image processing model is trained based on the initial image processing model, and the training of the initial image processing model can be end-to-end. A sample set with labels can be constructed, and supervised training can be performed on the initial image processing model based on the sample set.

[0148] The evaluation indicators of the key points obtained through the initial heat map include the classification accuracy and position accuracy of the key points. The classification accuracy indicator evaluates whether the key points are correctly identified as moving targets, usually measured by calculating the degree of match between the detected key points and the true key points. The position accuracy evaluates the deviation between the detected key points and the true position of the key points, usually using Euclidean distance or other distance metrics to calculate the positioning accuracy of the key points. These evaluation indicators can comprehensively evaluate the quality of the key points generated by the initial heat map, thereby ensuring its accuracy and reliability in target detection and positioning.

[0149] Specifically, the image samples to be analyzed are input into the initial image processing model to obtain target detection result samples and initial heat map samples; the image samples to be analyzed include moving target samples, and the moving target samples carry: position frame information, category information, position information of key point samples, and category information of the key point samples; the target detection result samples include predicted position frame information and predicted category information of the moving target samples; the initial heat map samples include predicted position information and predicted category information of the key point samples; according to the position frame information and category information of the moving target samples, and the target detection result samples, a target detection loss function is established; according to the position information and category information of the key point samples, and the predicted position information and predicted category information of the key point samples, a heat map loss function is established; based on the position information and category information of the key point samples, and the predicted position information and predicted category information of the key point samples, the initial image processing model is trained to obtain the trained image processing model.

[0150] An object detection loss function is established according to the difference between the position frame information of the moving object sample and the predicted position frame information in the object detection result sample, and according to the difference between the category information of the moving object sample and the predicted category information in the object detection result sample.

[0151] A heat map loss function is established based on the difference between the position information of the key point samples and the predicted position information of the key point samples, and the difference between the category information of the key point samples and the predicted category information of the key point samples.

[0152] A joint training loss function is determined according to the target detection loss function and the heat map loss function, and the initial image processing model is trained based on the joint training loss function. When the joint training loss function converges, the trained image processing model is obtained.

[0153] Figure 3 It is a flowchart for determining motion intention in the embodiment of the present disclosure. Target detection can be performed on the image to be analyzed first, and then the image to be analyzed can be divided into multiple image areas, an initial heat map is generated according to each image area, and the size of the initial heat map is adjusted according to the distance of the moving target corresponding to the initial heat map to obtain an adaptive heat map. The direction angle of the moving target is calculated based on the adaptive heat map. An evaluation index can be determined based on the target detection result and the initial heat map, and the evaluation index characterizes the accuracy of the target detection result and the accuracy of the key points in the initial heat map. Target tracking can be performed based on the key points to determine the driving intention of the moving target.

[0154] The technical solution of the disclosed embodiment adopts an adaptive heat map to detect the key points of the moving target, and dynamically adjusts the size of the heat map according to the distance between the object and the camera. The closer the object is, the larger the heat map is; the farther the object is, the smaller the heat map is. Such an adjustment can better adapt to objects at different distances, so that the position information of the object can be more accurately represented on the heat map, thereby improving the accuracy of key point detection, especially when the distance to the vehicle is very close, the direction angle can be more accurately judged. By accurately detecting the key points of the moving target, the technical solution of the disclosed embodiment improves the accurate judgment of the direction angle of the dynamic moving target. It avoids the accuracy problem of target detection at close distances, overcomes the shortcomings of the traditional method based on key points, and improves the accuracy of judging the driving intention. And adapt to complex scenes, due to the use of adaptive heat maps, the method can better adapt to complex scenes and reduce the degree of influence by factors such as occlusion and noise. It is more robust in actual road conditions and improves the applicability of the system in different environments.

[0155] It should be noted that, for the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present disclosure.

[0156] Figure 4 is a block diagram of a device for determining motion intention shown in an embodiment of the present disclosure, referring to Figure 3 The device includes an acquisition module, a detection module, a generation module, a matching module, an adjustment module and an analysis module, wherein:

[0157] An acquisition module, used for acquiring an image to be analyzed; the image to be analyzed includes at least one moving target;

[0158] A detection module, used to perform target detection on the image to be analyzed to obtain a target detection result; the target detection result includes: a predicted position frame of the moving target and a predicted category of the moving target;

[0159] A generating module, configured to generate at least one initial thermal map according to the image to be analyzed; the initial thermal map includes key point information of the moving target;

[0160] A matching module, used to match the target detection result with the initial heat map to obtain a moving target corresponding to each of the initial heat maps;

[0161] An adjustment module, used to adjust the size of the initial heat map corresponding to each of the moving targets according to the target detection result to obtain an adaptive heat map; the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map;

[0162] The analysis module is used to analyze the key points of the moving targets in the image to be analyzed according to the adaptive heat map corresponding to each moving target, so as to obtain the movement intention of each moving target.

[0163] Optionally, the analysis module is specifically used to perform:

[0164] According to the adaptive heat map corresponding to each of the moving targets, extracting a local image corresponding to each of the adaptive heat maps from the image to be analyzed, and adjusting the size of the local image to the size of the adaptive heat map corresponding to the local image;

[0165] According to the adaptive heat map corresponding to the local image, weighted processing is performed on the local image to obtain a weighted image;

[0166] According to the key points of the moving target in the weighted image, the direction angle of the moving target is calculated to obtain the moving direction intention of the moving target.

[0167] Optionally, the adjustment module is specifically used to execute:

[0168] According to the predicted category of the moving target, obtaining a correlation between the size of the position box of each moving target of the predicted category and the distance of the moving target;

[0169] Determining the distance of the moving target according to the association relationship and the size of the predicted position box of the moving target;

[0170] According to the distance of the moving target, the size of the initial heat map is adjusted to obtain the adaptive heat map.

[0171] Optionally, the generating module is specifically used to execute:

[0172] Acquire a feature map of the image to be analyzed;

[0173] Performing sliding convolution on the feature map to obtain a response value of each point of the feature map;

[0174] According to the response values ​​of each point of the feature map, the feature map is divided into at least one feature sub-map, and the feature sub-map containing the key point is determined as a local feature map;

[0175] Determine the point with the largest response value in the local feature map as the key point of the moving object;

[0176] Generate a response map according to the response value of the local feature map, and activate each of the response maps respectively to obtain a heat map to be filtered;

[0177] According to the distance between each point in the local feature map and the key point in the local feature map, Gaussian filtering is performed on the thermogram to be filtered to obtain the initial thermogram.

[0178] Optionally, when the image to be analyzed includes a plurality of continuous images, the analysis module is specifically configured to execute:

[0179] Tracking a moving target in each of the images to be analyzed;

[0180] For each of the moving targets, the motion attributes of the moving target are determined according to the adaptive thermal map corresponding to the moving target in each of the images to be analyzed and the key points of the moving target in the images to be analyzed; the motion attributes include: direction angle, speed, acceleration, historical motion trajectory and predicted motion trajectory;

[0181] The movement intention of the movement target is determined according to the movement attribute of the movement target.

[0182] Optionally, tracking a moving target in each of the images to be analyzed includes:

[0183] Establishing identifiers for key points of the moving target in a target image, and recording position information and feature information of the key points; the target image is: the first image among the multiple images to be analyzed;

[0184] In the event that the key point is lost in any of the images to be analyzed, predicting the position of the key point;

[0185] Matching and filtering the same key points in each of the images to be analyzed to obtain key point trajectories;

[0186] According to the key point trajectory, the position of the moving target in each of the images to be analyzed is determined.

[0187] Optionally, the detection module is specifically used to perform:

[0188] Acquire a feature map of the image to be analyzed;

[0189] Inputting the feature map into a region proposal network to obtain at least one candidate region, and determining a probability that the candidate region is the moving target;

[0190] Predicting the candidate area to obtain a plurality of initial position frames, probabilities of the initial position frames containing the moving target, and predicted categories of the moving target in the initial position frames;

[0191] According to the probability that the moving target is contained in the initial position frame, non-maximum suppression is performed on the initial position frame whose overlap ratio exceeds an overlap threshold to obtain the predicted position frame.

[0192] Optionally, the detection module is specifically used to perform:

[0193] Inputting the image to be analyzed into the target detection network of the image processing model to obtain the target detection result;

[0194] The generation module is specifically used to execute:

[0195] Inputting the image to be analyzed into the heat map generation network of the image processing model to obtain the initial heat map;

[0196] The training steps of the image processing model include:

[0197] Inputting the image samples to be analyzed into the initial image processing model, obtaining target detection result samples and initial heat map samples; the image samples to be analyzed include moving target samples, and the moving target samples carry: position frame information, category information, position information of key point samples, and category information of the key point samples; the target detection result samples include predicted position frame information and predicted category information of the moving target samples; the initial heat map samples include predicted position information and predicted category information of the key point samples;

[0198] Establishing a target detection loss function according to the position frame information and category information of the moving target sample and the target detection result sample;

[0199] Establishing a heat map loss function according to the position information and category information of the key point samples, and the predicted position information and predicted category information of the key point samples;

[0200] Based on the target detection loss function and the heat map loss function, the initial image processing model is trained to obtain the trained image processing model.

[0201] It should be noted that the device embodiment is similar to the method embodiment, so the description is relatively simple, and the relevant parts can be referred to the method embodiment.

[0202] The present disclosure also provides an electronic device, referring to Figure 5 , Figure 5 Schematic diagram of an electronic device according to an embodiment of the present disclosure. Figure 5As shown, the electronic device 100 includes: a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus communication. A computer program is stored in the memory 110. The computer program can be run on the processor 120 to implement the steps in the method for determining the motion intention disclosed in the embodiment of the present disclosure.

[0203] The embodiment of the present disclosure also provides a non-volatile readable storage medium. When the instructions in the non-volatile readable storage medium are executed by the processor of an electronic device, the electronic device can execute the steps in the method for determining motion intention as disclosed in the embodiment of the present disclosure.

[0204] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0205] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, devices or computer program products. Therefore, the embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0206] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0207] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0208] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0209] Although some embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the present disclosure.

[0210] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0211] The above is a detailed introduction to a method, device and electronic device for determining motion intention provided by the present disclosure. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core idea. At the same time, for a person skilled in the art, according to the idea of ​​the present disclosure, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

Claims

1. A method for determining movement intention, characterized in that: The method comprises: Acquire an image to be analyzed; the image to be analyzed includes at least one moving target; Performing target detection on the image to be analyzed to obtain a target detection result; the target detection result includes: a predicted position frame of the moving target and a predicted category of the moving target; Generate at least one initial thermal map according to the image to be analyzed; the initial thermal map includes key point information of the moving target; Matching the target detection result with the initial heat map to obtain a moving target corresponding to each initial heat map; According to the target detection result, the size of the initial heat map corresponding to each of the moving targets is adjusted to obtain an adaptive heat map; the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map; Analyzing the key points of the moving targets in the image to be analyzed according to the adaptive heat map corresponding to each moving target, and obtaining the movement intention of each moving target; The step of analyzing the key points of the moving targets in the image to be analyzed according to the adaptive heat map corresponding to each moving target to obtain the moving intention of each moving target includes: According to the adaptive heat map corresponding to each of the moving targets, extracting a local image corresponding to each of the adaptive heat maps from the image to be analyzed, and adjusting the size of the local image to the size of the adaptive heat map corresponding to the local image; According to the adaptive heat map corresponding to the local image, weighted processing is performed on the local image to obtain a weighted image; According to the key points of the moving target in the weighted image, the direction angle of the moving target is calculated to obtain the moving direction intention of the moving target.

2. The method according to claim 1, characterized in that The step of adjusting the size of the initial heat map corresponding to each of the moving targets according to the target detection result to obtain an adaptive heat map includes: According to the predicted category of the moving target, obtaining a correlation between the size of the position box of each moving target of the predicted category and the distance of the moving target; Determining the distance of the moving target according to the association relationship and the size of the predicted position box of the moving target; According to the distance of the moving target, the size of the initial heat map is adjusted to obtain the adaptive heat map.

3. The method according to claim 1, characterized in that The step of generating at least one initial thermal map according to the image to be analyzed includes: Acquire a feature map of the image to be analyzed; Performing sliding convolution on the feature map to obtain a response value of each point of the feature map; According to the response values ​​of each point of the feature map, the feature map is divided into at least one feature sub-map, and the feature sub-map containing the key point is determined as a local feature map; Determine the point with the largest response value in the local feature map as the key point of the moving object; Generate a response map according to the response value of the local feature map, and activate each of the response maps respectively to obtain a heat map to be filtered; According to the distance between each point in the local feature map and the key point in the local feature map, Gaussian filtering is performed on the thermogram to be filtered to obtain the initial thermogram.

4. The method according to claim 1, characterized in that: In the case where the image to be analyzed includes a plurality of continuous images, analyzing the key points of the moving target in the image to be analyzed according to the adaptive heat map corresponding to each moving target to obtain the movement intention of each moving target includes: Tracking a moving target in each of the images to be analyzed; For each of the moving targets, the motion attributes of the moving target are determined according to the adaptive thermal map corresponding to the moving target in each of the images to be analyzed and the key points of the moving target in the images to be analyzed; the motion attributes include: direction angle, speed, acceleration, historical motion trajectory and predicted motion trajectory; The movement intention of the movement target is determined according to the movement attribute of the movement target.

5. The method according to claim 4, characterized in that Tracking the moving target in each of the images to be analyzed includes: Establishing identifiers for key points of the moving target in a target image, and recording position information and feature information of the key points; the target image is: the first image among the multiple images to be analyzed; In the event that the key point is lost in any of the images to be analyzed, predicting the position of the key point; Matching and filtering the same key points in each of the images to be analyzed to obtain key point trajectories; According to the key point trajectory, the position of the moving target in each of the images to be analyzed is determined.

6. The method according to claim 1, characterized in that The performing target detection on the image to be analyzed to obtain a target detection result includes: Acquire a feature map of the image to be analyzed; Inputting the feature map into a region proposal network to obtain at least one candidate region, and determining a probability that the candidate region is the moving target; Predicting the candidate area to obtain a plurality of initial position frames, probabilities of the initial position frames containing the moving target, and predicted categories of the moving target in the initial position frames; According to the probability that the moving target is contained in the initial position frame, non-maximum suppression is performed on the initial position frame whose overlap ratio exceeds an overlap threshold to obtain the predicted position frame.

7. The method according to any one of claims 1 to 6, characterized in that: The performing target detection on the image to be analyzed to obtain a target detection result includes: Inputting the image to be analyzed into the target detection network of the image processing model to obtain the target detection result; The step of generating at least one initial thermal map according to the image to be analyzed includes: Inputting the image to be analyzed into the heat map generation network of the image processing model to obtain the initial heat map; The training steps of the image processing model include: Inputting the image samples to be analyzed into the initial image processing model, obtaining target detection result samples and initial heat map samples; the image samples to be analyzed include moving target samples, and the moving target samples carry: position frame information, category information, position information of key point samples, and category information of the key point samples; the target detection result samples include predicted position frame information and predicted category information of the moving target samples; the initial heat map samples include predicted position information and predicted category information of the key point samples; Establishing a target detection loss function according to the position frame information and category information of the moving target sample and the target detection result sample; Establishing a heat map loss function according to the position information and category information of the key point samples, and the predicted position information and predicted category information of the key point samples; Based on the target detection loss function and the heat map loss function, the initial image processing model is trained to obtain the trained image processing model.

8. A device for determining movement intention, characterized in that: The device comprises: An acquisition module, used for acquiring an image to be analyzed; the image to be analyzed includes at least one moving target; A detection module, used to perform target detection on the image to be analyzed to obtain a target detection result; the target detection result includes: a predicted position frame of the moving target and a predicted category of the moving target; A generating module, configured to generate at least one initial thermal map according to the image to be analyzed; the initial thermal map includes key point information of the moving target; A matching module, used to match the target detection result with the initial heat map to obtain a moving target corresponding to each of the initial heat maps; An adjustment module, used to adjust the size of the initial heat map corresponding to each of the moving targets according to the target detection result to obtain an adaptive heat map; the size of the adaptive heat map is inversely proportional to the distance of the moving target corresponding to the adaptive heat map; An analysis module, configured to analyze key points of the moving targets in the image to be analyzed according to an adaptive heat map corresponding to each moving target, so as to obtain a movement intention of each moving target; The analysis module is specifically used to perform: According to the adaptive heat map corresponding to each of the moving targets, extracting a local image corresponding to each of the adaptive heat maps from the image to be analyzed, and adjusting the size of the local image to the size of the adaptive heat map corresponding to the local image; According to the adaptive heat map corresponding to the local image, weighted processing is performed on the local image to obtain a weighted image; According to the key points of the moving target in the weighted image, the direction angle of the moving target is calculated to obtain the moving direction intention of the moving target.

9. An electronic device, characterized in that: include: processor; A memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method for determining the motion intention as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Moving target detection method and device, electronic equipment and storage medium

    CN113743219A

  • Face key point detection method and device, equipment and storage medium

    CN117649689A