Behavior prediction method and device, equipment, vehicle, storage medium and chip
By introducing global and local feature processing networks into the behavior prediction model, and carefully processing user behavior characteristics, the problem of low behavior prediction accuracy in the prior art is solved, and higher behavior prediction accuracy is achieved.
Patent Information
- Application Number
- CN202311696540.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the behavior prediction scheme has low accuracy and cannot effectively and meticulously handle the characteristics of user behavior.
A behavior prediction model including a global processing network, a local processing network and a behavior classification network is adopted to perform global and local feature processing on the pending video of the target object to generate predicted behavior.
Through global and local feature processing, the accuracy and accuracy of behavior prediction are significantly improved, and the roughness problem of traditional models in behavior feature processing is solved.
Smart Images

Figure CN120148097A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of big data processing, and particularly relates to a behavior prediction method, apparatus, device, vehicle, storage medium, and chip. Background Art
[0002] Behavior prediction is a technology for predicting the behavior of an object such as a user based on behavior prediction information (such as current environmental data, attribute information of the behavior execution object, etc.). Among them, user behavior prediction technology is widely applied in fields such as autonomous driving, security detection, and abnormal behavior detection in shopping malls.
[0003] In related technologies, traditional models (such as regression prediction models) are usually used to predict user behavior. However, traditional models such as regression prediction handle the features of user behavior rather roughly, resulting in a low accuracy of the behavior prediction scheme. Summary of the Invention
[0004] To overcome the problems existing in related technologies, the present disclosure provides a behavior prediction method, apparatus, device, vehicle, storage medium, and chip to solve the technical problem of low accuracy of the behavior prediction scheme existing in the above-mentioned related technologies.
[0005] According to the first aspect of the embodiments of the present disclosure, a behavior prediction method is provided, including:
[0006] Obtain a video to be processed of a target object, where the video to be processed includes at least one sequence frame;
[0007] Call a behavior prediction model to perform behavior prediction on the video to be processed of the target object, and obtain the predicted behavior of the target object;
[0008] Wherein, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network. The global processing network is used to perform global feature processing based on the video to be processed of the target object to obtain global feature information of the target object in each sequence frame. The local processing network is used to perform local feature processing based on the video to be processed of the target object to obtain local feature information of the target object in each sequence frame. The behavior classification network is used to perform behavior classification processing based on the global feature information and the local feature information of the target object in each sequence frame to obtain the predicted behavior of the target object.
[0009] In some embodiments, the global processing network includes a preprocessing sub-network and a global sub-network. The performing global feature processing based on the video to be processed of the target object to obtain global feature information of the target object in each sequence frame includes:
[0010] Call the preprocessing sub-network to perform tracking calculations on the video to be processed of the target object, and obtain the tracking information of the target object in each of the sequence frames. The tracking information at least includes the position area occupied by the target object in the sequence frame;
[0011] Call the global sub-network to perform global masking and encoding processing on the position area occupied by the target object in each of the sequence frames, and obtain the global feature information.
[0012] In some embodiments, the local processing network includes an N-level pyramid sub-network. The sampling windows respectively adopted by the N-level pyramid sub-networks increase in order. N is a positive integer. The local feature processing based on the video to be processed of the target object to obtain the local feature information of the target object in each of the sequence frames includes:
[0013] Perform region interception of a preset scale based on the position area occupied by the target object in each of the sequence frames to obtain the intercepted region of the target object in each of the sequence frames;
[0014] Call the N-level pyramid sub-network to perform local processing and superposition on the intercepted region of the target object in each of the sequence frames respectively to obtain the local feature information. The local processing includes at least one of the following processes: sampling, encoding, and masking.
[0015] In some embodiments, when N is 3, the local processing network includes a first-level pyramid sub-network, a second-level pyramid sub-network, and a third-level pyramid sub-network. The calling the N-level pyramid sub-network to perform local processing and superposition on the intercepted region of the target object in each of the sequence frames respectively to obtain the local feature information includes:
[0016] Call the first-level pyramid sub-network to perform sampling and encoding processing on the intercepted region of the target object in each of the sequence frames by using a preset first sampling window, and obtain the first feature information of the target object in each of the sequence frames;
[0017] Call the second-level pyramid sub-network to perform random sampling, masking, and encoding processing on the intercepted region of the target object in each of the sequence frames by using a preset second sampling window, and obtain the second feature information of the target object in each of the sequence frames;
[0018] Call the third-level pyramid sub-network to perform random sampling, masking, and encoding processing on the intercepted region of the target object in each of the sequence frames by using a preset third sampling window, and obtain the third feature information of the target object in each of the sequence frames;
[0019] Superimpose the first feature information, the second feature information, and the third feature information of the target object in each of the sequence frames to obtain the local feature information.
[0020] In some embodiments, the behavior classification network includes a decoding sub-network. The behavior classification process based on the global feature information and the local feature information of the target object in each of the sequence frames to obtain the predicted behavior of the target object includes:
[0021] Perform a connection process on the global feature information and the local feature information of the target object in each of the sequence frames to obtain the target feature information of the target object in each of the sequence frames;
[0022] Call the decoding sub-network to perform decoding and classification processing on the target feature information of the target object in each of the sequence frames to obtain the predicted behavior of the target object.
[0023] In some embodiments, the method further includes:
[0024] Obtain training samples, where the training samples include the training videos of the training objects and the labeled true behaviors;
[0025] Input the training videos of the training objects in the training samples into the behavior prediction model to be trained, and calculate the predicted behaviors of the training objects;
[0026] Based on the predicted behaviors and the true behaviors of the training objects, adjust the parameters of the behavior prediction model to be trained, so that when the loss value between the predicted behaviors and the true behaviors of the training objects is less than or equal to a preset value, obtain the trained behavior prediction model.
[0027] According to the second aspect of the embodiments of the present disclosure, a behavior prediction device is provided, including:
[0028] An acquisition module, configured to acquire a video to be processed of a target object, where the video to be processed includes at least one sequence frame;
[0029] A processing module, configured to call a behavior prediction model to perform behavior prediction on the video to be processed of the target object to obtain the predicted behavior of the target object;
[0030] Among them, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network. The global processing network is used to perform global feature processing on the video to be processed of the target object to obtain the global feature information of the target object in each sequence frame. The local processing network is used to perform local feature processing on the video to be processed of the target object to obtain the local feature information of the target object in each sequence frame. The behavior classification network is used to perform behavior classification processing based on the global feature information and the local feature information of the target object in each sequence frame to obtain the predicted behavior of the target object.
[0031] For the content not introduced or described in the embodiments of the present disclosure, reference may be made to the relevant introductions in the foregoing method embodiments correspondingly, and the embodiments of the present disclosure do not make any limitations.
[0032] According to a third aspect of the embodiments of the present disclosure, a terminal device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the executable instructions to implement the steps of the above-mentioned behavior prediction method.
[0033] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the behavior prediction method provided in the first aspect of the present disclosure are implemented.
[0034] According to a fifth aspect of the embodiments of the present disclosure, a chip is provided, including: a processor and an interface; the processor is used to read instructions to execute the steps of the above-mentioned behavior prediction method.
[0035] According to a sixth aspect of the embodiments of the present disclosure, a vehicle is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the executable instructions to implement the steps of the above-mentioned behavior prediction method.
[0036] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: The terminal device can obtain the video to be processed of the target object, and the video to be processed includes at least one sequence of frames; call the behavior prediction model to perform behavior prediction on the video to be processed of the target object to obtain the predicted behavior of the target object; wherein, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network. The global processing network is used to perform global feature processing on the video to be processed of the target object to obtain the global feature information of the target object in each of the sequence of frames. The local processing network is used to perform local feature processing on the video to be processed of the target object to obtain the local feature information of the target object in each of the sequence of frames. The behavior classification network is used to perform behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence of frames to obtain the predicted behavior of the target object. It can be seen that the terminal device can perform global and local feature processing on the video to be processed of the target object, so as to classify and predict the predicted behavior of the target object. Compared with the above related technologies, it can perform behavior feature processing more carefully or precisely, which is beneficial to improving the accuracy or precision of subsequent behavior prediction. At the same time, it can also solve the technical problems such as the low accuracy of the behavior prediction scheme in the above related technologies.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0039] Figure 1 is a flowchart of a behavior prediction method shown according to an exemplary embodiment.
[0040] Figure 2 is a schematic structural diagram of a behavior prediction model shown according to an exemplary embodiment.
[0041] Figure 3 is a schematic structural diagram of another behavior prediction model shown according to an exemplary embodiment.
[0042] Figure 4 is a processing flowchart of a behavior prediction model shown according to an exemplary embodiment.
[0043] Figure 5 is a schematic structural diagram of a behavior prediction device shown according to an exemplary embodiment.
[0044] Figure 6It is a schematic structural diagram of a terminal device shown according to an exemplary embodiment.
[0045] Figure 7 It is a schematic structural diagram of a chip shown according to an exemplary embodiment. Detailed implementation manners
[0046] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0047] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.
[0048] Large models refer to deep learning models with a very large number of parameters, such as Bert, GPT-3, etc.; these models have good performance because they can learn more features and patterns, thereby improving the accuracy and generalization ability of the models. Large models have a wider range of applications. For example, they can be applied to tasks or fields such as natural language processing, computer vision, and speech recognition. In addition, large models have higher efficiency, and can improve the efficiency of training and inference through technologies such as parallel computing and distributed training, thereby accelerating the training and inference processes of the models. Finally, large models have better interpretability, and they can help people understand the decision-making process and internal mechanisms of the models through visualization and interpretation technologies, thereby improving the interpretability and credibility of the models.
[0049] Behavior prediction (also known as behavior detection) has a wide range of applications in fields such as security, abnormal behavior detection in shopping malls (such as supermarkets), and autonomous driving. However, due to the large space occupied by videos themselves and the high demand for computing resources, it has led to the development of large models in this direction. Based on this, the present disclosure proposes a behavior prediction method, device, equipment, vehicle, storage medium, and chip, which can greatly reduce the computing power of video computing and improve the accuracy or precision of target object behavior prediction.
[0050] Please refer to Figure 1 It is a schematic flowchart of a behavior prediction method shown according to an exemplary embodiment. As Figure 1The method shown above can be applied to a terminal device, which can be, for example, a mobile phone, a tablet computer (Pad), a computer with wireless transceiver function, or a wireless terminal applied to scenarios such as virtual reality (VR), augmented reality (AR), industrial control, self-driving (also known as autonomous driving), remote medical, smart grid, transportation safety, smart city, and smart home. For example, it can be applied to a vehicle or a car in the self-driving scenario. The method may include the following implementation steps:
[0051] S101. Obtain a video to be processed of a target object, where the video to be processed includes at least one sequence of frames.
[0052] The video to be processed in the present disclosure refers to a video obtained by photographing a target object, which may include at least one continuous image frame, also known as a sequence of frames. The present disclosure does not limit the number of the above-mentioned target objects, which may be one or more. For the convenience of describing the present disclosure, the following content will be introduced taking one target object as an example, but it does not constitute a limitation. The above-mentioned target object may include, but is not limited to, pedestrians, users, puppies, kittens, or other custom objects, etc.
[0053] The present disclosure does not limit the implementation manner of obtaining the video to be processed. For example, it can be obtained by photographing the target object with a camera device, or obtained from other devices (such as a server or other terminals) through a network, etc.
[0054] S102. Invoke a behavior prediction model to perform behavior prediction on the video to be processed of the target object, and obtain the predicted behavior of the target object; wherein, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network. The global processing network is used to perform global feature processing based on the video to be processed of the target object to obtain the global feature information of the target object in each sequence of frames. The local processing network is used to perform local feature processing based on the video to be processed of the target object to obtain the local feature information of the target object in each sequence of frames. The behavior classification network is used to perform decoding and classification processing based on the global feature information and the local feature information of the target object in each sequence of frames to obtain the predicted behavior of the target object.
[0055] The above-mentioned behavior prediction model of the present disclosure is a pre-trained model installed in a terminal device, which is used to predict the behavior of a target object. The model includes, but is not limited to, for example, a convolutional neural network model, a deconvolution network model, a generative adversarial network model, a recurrent neural network model, or other custom large models or deep learning models. The present disclosure does not limit the internal structure of the above-mentioned behavior prediction model. For example, please refer to Figure 2 is a schematic diagram of the internal structure of a behavior prediction model shown according to an exemplary embodiment. As Figure 2 shown, the behavior prediction model may include: a global processing network 100, a local processing network 200, and a behavior classification network 300. Among them, the above-mentioned global processing network 100 is mainly used to perform global feature processing on the video to be processed of the target object, and obtain the global feature information of the target object in each sequence frame. The above-mentioned local processing network 200 is mainly used to perform local feature processing on the video to be processed of the target object, and obtain the local feature information of the target object in each sequence frame. The above-mentioned behavior classification network 300 is mainly used to perform behavior classification processing based on the global feature information and local feature information of the target object in each sequence frame, so as to predict the predicted behavior of the above-mentioned target object.
[0056] The present disclosure does not limit the internal structure and function implementation of the above-mentioned global processing network 100, local processing network 200, and behavior classification network 300 respectively. For example, please refer to Figure 3 is a schematic diagram of the internal structure of another behavior prediction model shown according to an exemplary embodiment. As Figure 3 shown in the model structure, the above-mentioned global processing network 100 may include a preprocessing sub-network 101 and a global sub-network 102. The above-mentioned local processing network 200 may include N-level (or N) pyramid sub-networks 201 connected in sequence. The sampling windows adopted by the N-level pyramid sub-networks 201 may be different. For example, usually the sampling windows adopted by the N-level pyramid sub-networks 201 increase in sequence. Among them, N is a positive integer customarily set according to actual situations. For example, N may be 3, etc. The above-mentioned behavior classification network 300 may include a decoding sub-network 301. Figure 3 The internal structure of the model shown is only an example. In actual applications, there may be more or fewer internal sub-networks according to actual situations, and the illustration does not constitute a limitation.
[0057] Please refer to Figure 4It is a processing flowchart of a behavior prediction model shown according to an exemplary embodiment. In a specific implementation, in the global processing network 100, the present disclosure may first call the above-mentioned preprocessing sub-network 101 to perform tracking calculations on the video to be processed of the target object, so as to obtain the tracking information of the target object in each sequence frame of the video. The tracking information may at least include the position area occupied by the target object in each sequence frame, such as the position information of the rectangular frame where the target object is located in the sequence frame (which can also be called bbox information, etc.). Optionally, the above-mentioned tracking information may further include information such as the identification (ID) information of the target object or other custom information. The present disclosure does not limit the specific implementation manners of the above-mentioned preprocessing sub-network and the above-mentioned tracking calculations respectively. For example, the present disclosure may use networks such as lightweight YOLOV8, BOT-SORT, and DEEP-SORT to perform object tracking calculations on each sequence frame in the video to be processed, so as to output the tracking information of the target object in the video, such as the detection box information of the target object in the sequence frame and the identification information of the target object. The present disclosure does not make too many limitations and detailed descriptions on this.
[0058] Furthermore, the present disclosure may call the global sub-network 102 to perform global masking and encoding processing on the position area occupied by the target object in each sequence frame. For example, specifically, it may first perform masking processing on the position area occupied by the target object in each sequence frame (i.e., the area where the detection box is located), and then perform encoding processing in sequence, so as to obtain the global feature information of the target object in each sequence frame. This is beneficial to reducing the data volume of video encoding, improving the computing power of video processing, and at the same time reducing the storage space size of the feature information.
[0059] In the local processing network 200, the present disclosure may first perform region cropping processing of a corresponding preset scale on each sequence frame based on the position area occupied by the target object obtained in each sequence frame, so as to obtain the cropped area of the target object in each sequence frame. The cropped area is usually larger than the corresponding position area. The above-mentioned preset scale is a cropping size or scale custom-set according to actual needs. For example, it can usually be 4 times the area size of the above-mentioned position area, etc. In this way, it can not only crop the position area occupied by the target object but also crop the corresponding background area occupied by the target object, which is convenient for the model to better learn and predict the behavior of the target object in the video, and is beneficial to improving the accuracy or accuracy rate of model behavior prediction / detection. It can be seen that the above-mentioned behavior prediction model of the present disclosure can perform user behavior prediction based on the temporal behavior of the target object at an accurate position obtained by the global processing network 100 and the relevant background information occupied by the target object obtained by the local processing network 200, thus being beneficial to improving the accuracy rate of behavior recognition / detection.
[0060] Furthermore, the present disclosure may call the above-mentioned N-level pyramid sub-network 201 to perform local processing and superposition processing on the intercepted region of the target object in each of the above-mentioned intercepted sequence frames, so as to obtain the local feature information of the target object in each sequence frame. The above-mentioned local processing may be determined according to actual situations. For example, it may include, but is not limited to, any one or a combination of the following processes: sampling, encoding, masking, or other custom processing, etc. See Figure 4 , taking N as 3 as an example, the above-mentioned local processing network 200 may include pyramid sub-networks of three scales, which may specifically include a first-level pyramid sub-network, a second-level pyramid sub-network, and a third-level pyramid sub-network. The sampling windows adopted by these three levels of pyramid sub-networks increase in sequence. For example, the sampling window size adopted by the above-mentioned first-level pyramid sub-network is 2×2, the sampling window size adopted by the above-mentioned second-level pyramid sub-network is 4×4, and the sampling window size adopted by the above-mentioned third-level pyramid sub-network is 6×6. Here, it is only for illustration and does not constitute a limitation, and it can be custom-adjusted according to actual situations. In specific implementation, the present disclosure may call the above-mentioned first-level pyramid sub-network to perform sampling and encoding processing on the intercepted region of the target object in each sequence frame by using a preset first sampling window (for example, it can be 2×2 here), so as to obtain the first feature information of the target object in each sequence frame. Call the above-mentioned second-level pyramid sub-network to perform random sampling, masking, and encoding processing on the intercepted region of the target object in each sequence frame by using a preset second sampling window (for example, it can be 4×4 here), so as to obtain the second feature information of the target object in each sequence frame. Specifically, the above-mentioned preset second sampling window may be used to perform random sampling on the intercepted region in each sequence frame, and then perform masking and encoding processing on the sampled pixels, etc., so as to correspondingly obtain the above-mentioned second feature information. The present disclosure may call the above-mentioned third-level pyramid sub-network to perform random sampling, masking, and encoding processing on the intercepted region of the target object in each sequence frame by using a preset third sampling window (for example, it can be 6×6 here), so as to obtain the third feature information of the target object in each sequence frame. It can be understood that both the above-mentioned second-level pyramid sub-network and the third-level pyramid sub-network adopt random sampling, and the purpose is to reduce the computational complexity of subsequent encoding, thereby improving the computing power of video processing. Finally, the present disclosure may perform superposition processing (also referred to as connection processing) on the obtained first feature information, second feature information, and third feature information of the target object in each sequence frame, so as to obtain the local feature information of the target object in each sequence frame.
[0061] In the behavior classification network 300, the present disclosure may first perform a connection process, that is, a stacking process, on the global feature information and local feature information of the above-mentioned target object in each sequence frame, so as to obtain the target feature information of the target object in each sequence frame, which may also be referred to as stacked feature information. Then, the decoding sub-network 301 is called to perform decoding and classification processing on the target feature information of the above-mentioned target object in each sequence frame, so as to output the predicted behavior of the above-mentioned target object. It can be understood that in practical applications, the above-mentioned decoding sub-network 301 may be a neural network structure for a decoder, which is usually used to convert the feature information generated by the encoding process (encoder) into a target output. Here, the final predicted behavior can be output, so as to achieve the ability to understand the input and generate an output. In image processing, a decoder is usually used to generate images, such as image generation, image restoration, etc. tasks, while in this embodiment, it is used to output the classification of different behaviors, that is, to output the final predicted behavior.
[0062] For example, taking the autonomous driving application scenario as an example, the present disclosure may collect videos of road pedestrians, and then call the above-mentioned behavior prediction model to perform behavior detection on the videos of road pedestrians, so as to obtain the predicted behavior of each pedestrian in the video. The predicted behavior here may include, but is not limited to, for example, pedestrians walking, sweeping the floor, dancing, sitting, standing, or other behavior classifications for describing pedestrians. For the specific implementation of how the above-mentioned behavior prediction model performs behavior detection, reference may be made to the relevant introduction in the foregoing embodiments, which will not be elaborated here.
[0063] In some alternative embodiments, the above-mentioned behavior prediction model needs to be trained before it is called. Among them, the training process of the above-mentioned behavior prediction model can be executed in a training device, which includes but is not limited to, for example, the above-mentioned terminal device, server, or other devices with model training capabilities, etc. The present disclosure does not make excessive limitations in this regard. In the embodiments of the present disclosure, supervised model learning and training can be performed on the behavior prediction model to be trained. In the specific implementation process, the present disclosure can obtain training samples, which include the training videos of the training objects and the true behaviors labeled for the training objects. Among them, the above-mentioned training videos can refer to the videos obtained by shooting the training objects, and the videos include a series of consecutive image frames, which can also be called sequence frames. The present disclosure does not limit the respective quantities of the above-mentioned training videos and training objects, and they can be determined according to the actual situation. For example, when there are multiple training videos and multiple training objects respectively, it means that the present disclosure can perform supervised model learning and training on different training objects in different training videos. When there is one training video and multiple training objects, it means that the present disclosure can perform supervised model learning and training on different training objects in the same training video. That is, the present disclosure can process different training objects in the same video and different training objects in different videos, such as encoding and supervised model parameter learning and training, etc.; this can improve the accuracy or precision of model training and help improve the accuracy of subsequent behavior prediction / classification.
[0064] After obtaining the above-mentioned training samples, the present disclosure can input the training videos of the training objects in the above-mentioned training samples into the behavior prediction model to be trained for behavior prediction, calculate and output the predicted behaviors of the above-mentioned training objects. The specific implementation of the above-mentioned behavior prediction is not limited and elaborated here too much, such as encoding, decoding, and behavior classification processing, etc. Then, based on the predicted behaviors of the training objects and the true behaviors labeled for the training objects, the model parameters of the above-mentioned behavior prediction model to be trained are continuously adjusted, so that the model is optimized in the correct target classification direction. Specifically, when the loss value between the predicted behavior of the above-mentioned training object and its true behavior is small (for example, it can be less than or equal to a preset value), the finally trained above-mentioned behavior prediction model is output. The preset value can be a value custom-set according to the actual situation, such as an empirical value set according to user experience, or a statistical value calculated based on a series of experimental data, etc. Regarding the specific implementation details of how the present disclosure trains the above-mentioned behavior prediction model based on training samples, the present disclosure does not expand and describe and limit too much here.
[0065] By implementing the embodiments of the present disclosure, a terminal device can obtain a video to be processed of a target object, where the video to be processed includes at least one sequence of frames; call a behavior prediction model to perform behavior prediction on the video to be processed of the target object to obtain the predicted behavior of the target object; wherein, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network, the global processing network is configured to perform global feature processing based on the video to be processed of the target object to obtain global feature information of the target object in each of the sequence of frames, the local processing network is configured to perform local feature processing based on the video to be processed of the target object to obtain local feature information of the target object in each of the sequence of frames, and the behavior classification network is configured to perform behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence of frames to obtain the predicted behavior of the target object. It can be seen that the terminal device can perform global and local feature processing on the video to be processed of the target object, so as to classify and predict the predicted behavior of the target object. Compared with the above related technologies, it can perform behavior feature processing more carefully or precisely, which is beneficial to improving the accuracy or precision of subsequent behavior prediction. At the same time, it can also solve the technical problems such as the low accuracy of the behavior prediction scheme in the above related technologies.
[0066] Based on the foregoing embodiments, please refer to Figure 5 is a schematic structural diagram of a behavior prediction device shown according to an exemplary embodiment. As Figure 5 shown, the device can be applied to a terminal device, and the device can include an acquisition module 501 and a processing module 502. Wherein:
[0067] The acquisition module 501 is configured to acquire a video to be processed of a target object, where the video to be processed includes at least one sequence of frames;
[0068] The processing module 502 is configured to call a behavior prediction model to perform behavior prediction on the video to be processed of the target object to obtain the predicted behavior of the target object;
[0069] Wherein, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network, the global processing network is configured to perform global feature processing based on the video to be processed of the target object to obtain global feature information of the target object in each of the sequence of frames, the local processing network is configured to perform local feature processing based on the video to be processed of the target object to obtain local feature information of the target object in each of the sequence of frames, and the behavior classification network is configured to perform behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence of frames to obtain the predicted behavior of the target object.
[0070] In some embodiments, the global processing network includes a preprocessing sub-network and a global sub-network. The processing module 502 is configured to call the preprocessing sub-network to perform tracking calculations on the video to be processed of the target object, obtain the tracking information of the target object in each of the sequence frames, and the tracking information at least includes the position area occupied by the target object in the sequence frame; call the global sub-network to perform global masking and encoding processing on the position area occupied by the target object in each of the sequence frames to obtain the global feature information.
[0071] In some embodiments, the local processing network includes an N-level pyramid sub-network, and the sampling windows respectively adopted by the N-level pyramid sub-networks increase sequentially in order. N is a positive integer. The processing module 502 is configured to perform region interception of a preset scale based on the position area occupied by the target object in each of the sequence frames to obtain the intercepted area of the target object in each of the sequence frames; call the N-level pyramid sub-networks to respectively perform local processing and superposition on the intercepted area of the target object in each of the sequence frames to obtain the local feature information, and the local processing includes at least one of the following processes: sampling, encoding, and masking.
[0072] In some embodiments, when N is 3, the local processing network includes a first-level pyramid sub-network, a second-level pyramid sub-network, and a third-level pyramid sub-network. The processing module 502 is configured to call the first-level pyramid sub-network to perform sampling and encoding processing on the intercepted area of the target object in each of the sequence frames by using a preset first sampling window to obtain the first feature information of the target object in each of the sequence frames; call the second-level pyramid sub-network to perform random sampling, masking, and encoding processing on the intercepted area of the target object in each of the sequence frames by using a preset second sampling window to obtain the second feature information of the target object in each of the sequence frames; call the third-level pyramid sub-network to perform random sampling, masking, and encoding processing on the intercepted area of the target object in each of the sequence frames by using a preset third sampling window to obtain the third feature information of the target object in each of the sequence frames; superimpose the first feature information, the second feature information, and the third feature information of the target object in each of the sequence frames to obtain the local feature information.
[0073] In some embodiments, the behavior classification network includes a decoding sub-network. The processing module 502 is configured to perform connection processing on the global feature information and the local feature information of the target object in each of the sequence frames to obtain the target feature information of the target object in each of the sequence frames; call the decoding sub-network to perform decoding and classification processing on the target feature information of the target object in each of the sequence frames to obtain the predicted behavior of the target object.
[0074] In some embodiments, the processing module 502 is further configured to obtain training samples, where the training samples include training videos of a training object and labeled true behaviors; input the training videos of the training object in the training samples into the behavior prediction model to be trained, and calculate the predicted behaviors of the training object; and adjust the parameters of the behavior prediction model to be trained based on the predicted behaviors and the true behaviors of the training object, so as to obtain the trained behavior prediction model when the loss value between the predicted behaviors and the true behaviors of the training object is less than or equal to a preset value.
[0075] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0076] By implementing the embodiments of the present disclosure, the above device can obtain a video to be processed of a target object, where the video to be processed includes at least one sequence of frames; call a behavior prediction model to perform behavior prediction on the video to be processed of the target object, and obtain the predicted behaviors of the target object; where the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network, the global processing network is configured to perform global feature processing on the video to be processed of the target object to obtain global feature information of the target object in each of the sequence of frames, the local processing network is configured to perform local feature processing on the video to be processed of the target object to obtain local feature information of the target object in each of the sequence of frames, and the behavior classification network is configured to perform behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence of frames to obtain the predicted behaviors of the target object. It can be seen that the terminal device can perform global and local feature processing on the video to be processed of the target object, so as to classify and predict the predicted behaviors of the target object. Compared with the above related technologies, it can perform behavior feature processing more carefully or precisely, which is beneficial to improving the accuracy or precision of subsequent behavior prediction. At the same time, it can also solve the technical problems such as the low accuracy of the behavior prediction scheme in the above related technologies.
[0077] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the behavior prediction method provided by the present disclosure are implemented.
[0078] Figure 6It is a schematic structural diagram of a terminal device shown according to an exemplary embodiment. For example, the terminal device 600 may be a vehicle (or an automobile), a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, or other terminal devices.
[0079] Referring to Figure 6 , the terminal device 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output interface 612, a sensor component 614, and a communication component 616.
[0080] The processing component 602 generally controls the overall operation of the device 600, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above-mentioned behavior prediction method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.
[0081] The memory 604 is configured to store various types of data to support the operation of the device 600. Examples of these data include instructions for any application or method operating on the device 600, contact data, phone book data, messages, pictures, videos, and the like. The memory 604 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0082] The power component 606 provides power to various components of the device 600. The power component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 600.
[0083] The multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the terminal device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0084] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive external audio signals when the device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.
[0085] The input / output interface 612 provides an interface between the processing component 602 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0086] The sensor component 614 includes one or more sensors for providing status assessments of various aspects of the device 600. For example, the sensor component 614 can detect the on / off state of the device 600, the relative positioning of components, such as the display and the keypad of the device 600. The sensor component 614 can also detect a change in the position of the device 600 or a component of the device 600, the presence or absence of user contact with the device 600, the orientation or acceleration / deceleration of the device 600, and the temperature change of the device 600. The sensor component 614 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 614 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 614 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0087] The communication component 616 is configured to facilitate communication between the device 600 and other devices in a wired or wireless manner. The device 600 may access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0088] In an exemplary embodiment, the device 600 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described behavior prediction method.
[0089] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the above instructions can be executed by a processor 620 of the device 600 to complete the above-described upper behavior prediction method. For example, the non-transitory computer-readable storage medium may be a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0090] In addition to being an independent electronic device, the above device can also be a part of an independent electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip. The integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), SOC (System on Chip), etc. The above integrated circuit or chip can be used to execute executable instructions (or code) to implement the above behavior prediction method. The executable instructions can be stored in the integrated circuit or chip, or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, a memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above behavior prediction method is implemented. Alternatively, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above behavior prediction method.
[0091] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above behavior prediction method when executed by the programmable device.
[0092] Please refer to Figure 7 is a schematic structural diagram of a chip shown according to an exemplary embodiment. As Figure 7 shown, the chip 700 includes a processor 701 and an interface 702. Optionally, a memory 703 can also be included. Among them, the number of processors 701 can be one or more, and the number of interfaces 702 can be multiple.
[0093] In one embodiment, for the case where the chip is used to implement the method embodiments of the present disclosure:
[0094] The interface 702 is used to receive or output signals;
[0095] The processor 701 is used to execute some or all of the content in the method embodiments of the behavior prediction.
[0096] Understandably, the processor in the embodiments of the present disclosure may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method embodiments may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0097] Understandably, the memory in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memories of the systems and methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0098] It should be noted here that the descriptions of the above storage medium, device, and chip embodiments are similar to those of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium, storage medium, and device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.
[0099] After considering the specification and practicing the present disclosure, those skilled in the art will readily conceive of other implementations of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0100] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A behavior prediction method, characterized in that, it includes: Obtain the video to be processed of the target object, where the video to be processed includes at least one sequence of frames; Call a behavior prediction model to perform behavior prediction on the video to be processed of the target object to obtain the predicted behavior of the target object; Wherein, the behavior prediction model includes a global processing network, a local processing network and a behavior classification network. The global processing network is used to perform global feature processing based on the video to be processed of the target object to obtain the global feature information of the target object in each of the sequence of frames. The local processing network is used to perform local feature processing based on the video to be processed of the target object to obtain the local feature information of the target object in each of the sequence of frames. The behavior classification network is used to perform behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence of frames to obtain the predicted behavior of the target object.
2. The method according to claim 1, characterized in that, The global processing network includes a preprocessing sub-network and a global sub-network. The performing global feature processing based on the video to be processed of the target object to obtain the global feature information of the target object in each of the sequence of frames includes: Call the preprocessing sub-network to perform tracking calculation on the video to be processed of the target object to obtain the tracking information of the target object in each of the sequence of frames, where the tracking information at least includes the position area occupied by the target object in the sequence of frames; Call the global sub-network to perform global masking and encoding processing on the position area occupied by the target object in each of the sequence of frames to obtain the global feature information.
3. The method according to claim 2, characterized in that, The local processing network includes an N-level pyramid sub-network, and the sampling windows adopted by the N-level pyramid sub-networks increase in sequence. N is a positive integer. The performing local feature processing based on the video to be processed of the target object to obtain the local feature information of the target object in each of the sequence of frames includes: Perform region interception of a preset scale based on the position area occupied by the target object in each of the sequence of frames to obtain the intercepted region of the target object in each of the sequence of frames; Call the N-level pyramid sub-networks to perform local processing and superposition on the intercepted regions of the target object in each of the sequence of frames respectively to obtain the local feature information, and the local processing includes at least one of the following processes: sampling, encoding and masking.
4. The method according to claim 3, characterized in that, When N is 3, the local processing network includes a first-level pyramid sub-network, a second-level pyramid sub-network and a third-level pyramid sub-network. The calling the N-level pyramid sub-networks to perform local processing and superposition on the intercepted regions of the target object in each of the sequence of frames respectively to obtain the local feature information includes: Call the first-level pyramid sub-network to sample and encode the intercepted area of the target object in each of the sequence frames using a preset first sampling window, to obtain the first feature information of the target object in each of the sequence frames; Call the second-level pyramid sub-network to randomly sample, mask, and encode the intercepted area of the target object in each of the sequence frames using a preset second sampling window, to obtain the second feature information of the target object in each of the sequence frames; Call the third-level pyramid sub-network to randomly sample, mask, and encode the intercepted area of the target object in each of the sequence frames using a preset third sampling window, to obtain the third feature information of the target object in each of the sequence frames; Superimpose the first feature information, the second feature information, and the third feature information of the target object in each of the sequence frames to obtain the local feature information.
5. The method according to claim 1, wherein, the behavior classification network includes a decoding sub-network, and the performing behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence frames to obtain the predicted behavior of the target object includes: Performing a connection process on the global feature information and the local feature information of the target object in each of the sequence frames to obtain the target feature information of the target object in each of the sequence frames; Call the decoding sub-network to perform decoding and classification processing on the target feature information of the target object in each of the sequence frames to obtain the predicted behavior of the target object.
6. The method according to any one of claims 1-5, wherein, the method further includes: Obtaining a training sample, where the training sample includes a training video of a training object and an annotated true behavior; Inputting the training video of the training object in the training sample into the behavior prediction model to be trained, and calculating to obtain the predicted behavior of the training object; Based on the predicted behavior and the true behavior of the training object, adjusting the parameters of the behavior prediction model to be trained, so that when the loss value between the predicted behavior and the true behavior of the training object is less than or equal to a preset value, obtaining the trained behavior prediction model.
7. A behavior prediction device, wherein, it includes: An acquisition module configured to acquire a video to be processed of a target object, where the video to be processed includes at least one sequence frame; A processing module configured to call a behavior prediction model to perform behavior prediction on the video to be processed of the target object to obtain the predicted behavior of the target object; Among them, the behavior prediction model includes a global processing network, a local processing network, and a behavior classification network. The global processing network is used to perform global feature processing on the video to be processed of the target object to obtain the global feature information of the target object in each of the sequence frames. The local processing network is used to perform local feature processing on the video to be processed of the target object to obtain the local feature information of the target object in each of the sequence frames. The behavior classification network is used to perform behavior classification processing based on the global feature information and the local feature information of the target object in each of the sequence frames to obtain the predicted behavior of the target object.
8. A terminal device Characterized in that It includes: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to execute the executable instructions to implement the steps of the method according to any one of claims 1 to 6.
9. A vehicle Characterized in that It includes: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to execute the executable instructions to implement the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, on which computer program instructions are stored Characterized in that When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
11. A chip Characterized in that It includes a processor and an interface; the processor is used to read instructions to execute the method according to any one of claims 1 to 6.