Behavior prediction model training method, behavior prediction method and system

By training and updating the behavior prediction model on edge computing devices, combining reward and punishment scores and hyperparameter tuning, the accuracy and security problems of the behavior prediction model in the prior art in complex scenarios are solved, and efficient and safe behavior prediction is achieved.

CN120259762APending Publication Date: 2025-07-04EAPIL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510365167.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing behavior prediction model is difficult to adapt to environmental changes and diversified behavioral patterns in complex and changeable monitoring scenarios, and is inaccurate and has prominent data transmission and security issues.

Method used

The behavior prediction model is trained on the edge computing device, and the training model prediction results are updated iteratively, combining reward and punishment scores and hyperparameter tuning algorithm to optimize the model strategy, reduce data transmission, and improve the accuracy and security of the model in complex scenarios.

Benefits of technology

It improves the accuracy and processing efficiency of the behavior prediction model in complex scenarios, reduces data transmission requirements, reduces network pressure and security risks, and ensures the stability and security of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259762A_ABST
    Figure CN120259762A_ABST
Patent Text Reader

Abstract

The invention provides a behavior prediction model training method, a behavior prediction method and a behavior prediction system, the behavior prediction model training method is applied to edge computing equipment, and the method comprises the following steps: training a basic model through sample data; the model obtained by training the basic model comprises a behavior prediction model; predicting the behavior of the target image through the trained model, and outputting a prediction result; wherein the target image is acquired through image acquisition equipment in a set area; updating the trained model according to the prediction result; and continuing to train the updated model through the new sample data until the updated model training reaches an iteration condition. In the model training process, the model is iteratively updated according to the prediction result of the trained model, so that the accuracy of model training can be improved, and the prediction accuracy of the behavior prediction model obtained through training is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image monitoring, and in particular, to a method for training a behavior prediction model, a behavior prediction method, and a system. Background Art

[0002] In today's society, the monitoring field plays a crucial role and is widely used in many important fields such as security, transportation, and urban management. With the continuous development of technology, in order to manage and respond to various situations more efficiently, the application of behavior prediction models in the monitoring field has received increasing attention. These models aim to predict the possible behaviors of individuals or groups by analyzing historical data and real-time information, so as to make corresponding decisions and arrangements in advance.

[0003] However, the monitoring scenarios in the real environment are extremely complex and variable, bringing huge challenges to behavior prediction models. The scenarios involved in the monitoring field usually contain a large number of uncertain factors, such as the diversity of people, the dynamic changes of the environment, the unpredictability of emergencies, etc. Most of the existing methods rely on limited training data and fixed feature extraction methods, and it is difficult to adapt to the real-time changing environment and diverse behavior patterns. Summary of the Invention

[0004] In view of this, the purpose of the embodiments of this application is to provide a method for training a behavior prediction model, a behavior prediction method, and a system, so as to improve the accuracy of the behavior prediction model in complex and variable scenarios.

[0005] In a first aspect, the embodiments of this application provide a method for training a behavior prediction model, which is applied to an edge computing device. The method includes: training a basic model with sample data; wherein, the model obtained by training the basic model includes a behavior prediction model; predicting the behavior of a target image through the trained model and outputting a prediction result; wherein, the target image is obtained by an image acquisition device in a set area; updating the trained model according to the prediction result; continuing to train the updated model with new sample data until the updated model training reaches the iteration condition.

[0006] In the above implementation process, by training the behavior prediction model on the edge computing device, and during the model training process, iteratively updating the model through the prediction results of the trained model, the accuracy of the behavior prediction model in complex and variable scenarios can be improved, and at the same time, data transmission can be reduced, and the processing efficiency and data security can be improved.

[0007] In one embodiment, updating the trained model according to the prediction result includes: determining a reward and punishment score based on the prediction result and the actual result; wherein, the reward and punishment score is determined according to the accuracy of the prediction result; determining an optimization strategy through the reward and punishment score and an adjustment strategy; and updating the prediction strategy of the trained model according to the optimization strategy.

[0008] In the above implementation process, the prediction result of the trained model is scored through the reward and punishment score, and then the corresponding optimization strategy is determined according to the reward and punishment score, and the prediction strategy of the trained model is updated, which can improve the accuracy of the prediction strategy in the model, and further improve the accuracy of the behavior prediction model.

[0009] In one embodiment, updating the trained model according to the prediction result includes: determining an evaluation index of the trained model according to multiple prediction results of the trained model for behavior prediction of the target image; and adjusting hyperparameters in the trained model according to the evaluation index and a hyperparameter tuning algorithm.

[0010] In the above implementation process, by adjusting the hyperparameters in the trained model according to the evaluation index and the hyperparameter tuning algorithm, the optimal parameter settings corresponding to the model can be determined, and the overall performance of the model can be improved.

[0011] In one embodiment, updating the trained model according to the prediction result includes: obtaining first model parameters of a trained related model; and adjusting second model parameters of the trained model through the first model parameters.

[0012] In the above implementation process, by adjusting the second model parameters of the trained model by using the first model parameters of the trained related model, fine-tuning of the model can be realized by using a small amount of data, the model convergence speed can be improved, the training efficiency can be improved, and at the same time, the dependence on data annotation can be reduced.

[0013] In a second aspect, an embodiment of the present application further provides a behavior prediction method applied to an edge computing device, and the method includes: inputting the target image into a behavior prediction model; wherein, the behavior prediction model is obtained according to the behavior prediction model training method in the first aspect or any possible implementation manner of the first aspect; and predicting the behavior corresponding to the target image through the behavior prediction model.

[0014] In the above implementation process, by using the behavior prediction model stored on the edge computing device to perform behavior prediction on the target image, a large amount of data transmission can be reduced, the requirement for network bandwidth can be reduced, network congestion can be effectively avoided, and the stability of behavior prediction in a complex network environment can be improved. Even in areas with poor network conditions, real-time monitoring and behavior prediction can be realized, and the accuracy of behavior prediction can be improved.

[0015] In one embodiment, predicting the behavior corresponding to the target image through the behavior prediction model includes: determining the historical data and current behavior data corresponding to the target image; inputting the historical data and the current behavior data into an autoregressive model to generate a prediction value; classifying through a softmax function and the prediction value pair to determine the probability value of each behavior; and determining the behavior with the largest probability value as the behavior corresponding to the target image.

[0016] In the above implementation process, when performing behavior prediction, determining the behavior with the largest probability value as the behavior corresponding to the target image can improve the accuracy of behavior prediction.

[0017] In a third aspect, an embodiment of the present application further provides a behavior prediction system, including: an image acquisition device and an edge computing device; the edge computing device is connected to the image acquisition device; the image acquisition device is configured to acquire a target image and transmit the target image to the edge computing device; the edge computing device is configured to predict the behavior corresponding to the target image through a behavior prediction model; wherein, the behavior prediction model is obtained according to the behavior prediction model training method in the first aspect, or any possible implementation manner of the first aspect, and the behavior prediction model is stored in the edge computing device.

[0018] In the above implementation process, by setting an edge computing device and an image acquisition device in the behavior prediction system, and connecting the edge computing device and the image acquisition device, the target image acquired by the image acquisition device can be directly transmitted to the edge computing device, and then the edge computing device can predict the behavior in the corresponding area of the target image according to the target image, reducing the data transmitted to the remote server, improving the efficiency of behavior monitoring, while reducing the network pressure and lowering the security risk during data transmission.

[0019] In one embodiment, it further includes: a server; the edge computing device is connected to the server; the edge computing device is further configured to transmit key feature data to the server; wherein, the key feature data is the main feature data for behavior prediction; the key feature data includes contour data, pose data, and position data.

[0020] In the above implementation process, by setting a server and connecting the server to the edge computing device, data that is difficult to process by the edge computing device can be transmitted to the server for further processing, improving the data processing ability of the behavior prediction system. And when performing data transmission, only transmitting the key feature data can reduce the data transmission volume, lower the network bandwidth pressure, and improve the data transmission efficiency and stability.

[0021] In one embodiment, the edge computing device is configured to adjust the operating power consumption by adjusting the operating frequency and voltage of the processor; wherein, the formula for calculating the operating power consumption is: P = f × C × V 2 ; where P is the operating power consumption, f is the operating frequency of the processor, C is the load capacitance, and V is the voltage of the processor.

[0022] In the above implementation process, by setting the processor of the edge processing device to be able to adjust its operating power according to the operating frequency and voltage, it is possible to improve the computing performance while reducing the power consumption of the edge processing device.

[0023] In one embodiment, the edge computing device includes a disk array; the edge computing device is configured to store the data for behavior prediction dispersedly on multiple independent disks in the disk array.

[0024] In the above implementation process, by using a disk array for data storage, when a certain disk fails, the data can be restored through other disks, which can prevent data loss while controlling costs.

[0025] Fourthly, an embodiment of the present application further provides a behavior prediction model training device, which is applied to an edge computing device and includes: a training module for training a basic model through sample data; wherein, the model obtained by training the basic model includes a behavior prediction model; a first prediction module for predicting the behavior of a target image through the trained model and outputting a prediction result; wherein, the target image is obtained by an image acquisition device in a set area; an update module for updating the trained model according to the prediction result; and the training module is further configured to continue training the updated model through new sample data until the updated model training reaches the iteration condition.

[0026] Fifthly, an embodiment of the present application further provides a behavior prediction device, which is applied to an edge computing device and includes: an input module for inputting the target image into a behavior prediction model; wherein, the behavior prediction model is obtained according to the behavior prediction model training method in the first aspect or any possible implementation manner of the first aspect; a second prediction module for predicting the behavior corresponding to the target image through the behavior prediction model.

[0027] Sixthly, an embodiment of the present application further provides an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, when the machine-readable instructions are executed by the processor, the steps of the methods in the first aspect, or any possible implementation manner of the first aspect, the second aspect, or any possible implementation manner of the second aspect are executed.

[0028] In a seventh aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method in the above first aspect, or any possible implementation manner of the first aspect, the second aspect, or any possible implementation manner of the second aspect.

[0029] To make the above objects, features, and advantages of the present application more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0031] Figure 1 A schematic diagram of the interaction of devices in the behavior prediction system provided by the embodiment of the present application;

[0032] Figure 2 A flowchart of the method for training a behavior prediction model provided by the embodiment of the present application;

[0033] Figure 3 A flowchart of the behavior prediction method provided by the embodiment of the present application;

[0034] Figure 4 A schematic diagram of the functional modules of the device for training a behavior prediction model provided by the embodiment of the present application;

[0035] Figure 5 A schematic diagram of the functional modules of the behavior prediction device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0037] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0038] In today's highly digitalized era, video surveillance has been widely integrated into all aspects of society. From urban security surveillance networks, intelligent transportation traffic condition monitoring, to commercial venue operation management and home security protection, all rely on the support of camera systems. However, traditional camera systems have many insurmountable drawbacks. Most of them can only achieve basic video acquisition functions, and image recognition is based on simple algorithms, such as common face recognition. The algorithms are single, and there is great room for improvement in accuracy and adaptability. In the face of complex and changeable real-world scenarios, the deficiencies of traditional systems are particularly obvious.

[0039] In the field of security surveillance, traditional cameras can only record the whole process of events and cannot effectively give early warnings before events occur. Take crowded and complex places such as large shopping malls and railway stations as examples. When potential dangerous behaviors such as abnormal gathering of people and possible conflicts occur, traditional camera systems often respond only after the event, resulting in the missed best prevention and handling opportunities and making it difficult to effectively protect the personal and property safety of the public. For example, when a thief is about to commit theft in a shopping mall, traditional cameras cannot detect abnormal behaviors in advance, and only record the pictures after the theft occurs, which is of little significance for recovering losses.

[0040] In intelligent transportation, traditional camera systems lack the ability to accurately predict abnormal behaviors of vehicles and pedestrians. At intersections, dangerous behaviors such as sudden illegal lane changes by vehicles and pedestrians running red lights occur frequently. Due to the lack of effective prediction ability, traditional cameras cannot make advance judgments, and it is difficult for traffic management departments to take measures in advance to avoid traffic accidents, which not only seriously affects traffic flow but also poses a great threat to the safety of pedestrians and vehicles. For example, during the morning and evening rush hours, illegal lane changes by vehicles are likely to cause chain rear-end accidents, and traditional camera systems cannot give early warnings, making traffic congestion and accident handling more difficult.

[0041] Existing camera systems generally adopt a mode of transmitting a large amount of raw video data to a cloud server for centralized processing in the data transmission and processing process. This mode exposes a series of serious problems in practical applications. On the one hand, the transmission of a large amount of high-definition video data requires extremely high network bandwidth. With the popularization of high-definition (1080p), ultra-high-definition (4K or even 8K) video surveillance, the amount of video data has increased exponentially. Taking a 1080p, 30fps video as an example, calculated according to each pixel occupying 24 bits (3 bytes), the data volume of a single frame is about 1920×1080×3≈6.22MB; the data volume per second of a 30fps video stream: 6.22MB×30≈186.6MB / s. Such a huge amount of data is extremely likely to cause network congestion in areas with poor network conditions or during peak network usage periods, resulting in data transmission delays or even interruptions, seriously affecting the real-time performance and stability of the monitoring system.

[0042] Moreover, data faces numerous security risks during transmission. Especially when it comes to personal privacy and important business information, data security issues cannot be ignored. During transmission, data may be stolen or tampered with by hackers, resulting in serious consequences. Take the smart home monitoring scenario as an example. Once the user's home life video data is leaked, it not only violates the user's privacy but may also be exploited by criminals, causing property losses and even threats to personal safety. In the business field, enterprise monitoring data may include trade secrets such as production processes and customer information. Once the data is tampered with or stolen, it will have a huge impact on the operation and development of the enterprise.

[0043] In view of this, an embodiment of the present application proposes a method for training a behavior prediction model. By training the behavior prediction model on an edge computing device and iteratively updating the model based on the prediction results of the trained model, the accuracy of the behavior prediction model in complex and variable scenarios can be improved. At the same time, data transmission can be reduced, and the processing efficiency and data security can be enhanced.

[0044] To facilitate the understanding of this embodiment, a behavior prediction system disclosed in the embodiment of the present application will be introduced in detail first.

[0045] As Figure 1 shown, it is a schematic diagram of the interaction between devices in the behavior prediction system provided by the embodiment of the present application, including: an image acquisition device 200 and an edge computing device 100.

[0046] Among them, the edge computing device 100 is connected to the image acquisition device 200.

[0047] Optionally, the edge computing device 100 communicates with one or more image acquisition devices 200 through a network.

[0048] The image acquisition device 200 here is used to acquire a target image and transmit the target image to the edge computing device 100. The image acquisition device 200 can be a camera, a sky network, a monitoring device, etc. One or more image acquisition devices 200 may be included in the behavior prediction system, and the type and quantity of the image acquisition device 200 can be selected according to the actual situation.

[0049] In one embodiment, the image acquisition device 200 can transmit the acquired target image to the edge computing device 100 in real time at a rate of 100 Mbps. After receiving the target image, the edge computing device 100 performs preliminary processing and analysis on the target image, and extracts key behavior features such as human actions, postures, and speeds.

[0050] Among them, the image acquisition device 200 can adopt a high-resolution image sensor. The pixel of the image acquisition device 200 is greater than or equal to 8 million.

[0051] The image acquisition device 200 can be used to capture fine image details in the monitoring area and can also form clear images in low-light conditions.

[0052] In one embodiment, the image acquisition device 200 has an ultra-wide-angle shooting ability, and its viewing angle range can reach 180 degrees. For example, in urban road monitoring, one image acquisition device 200 (such as a camera terminal) can cover multiple lanes and the surrounding sidewalk areas to improve the monitoring efficiency.

[0053] Among them, the viewing angle range can be determined by the following formula:

[0054] h = d × tan(θ);

[0055] Among them, d is the monitoring distance, θ is the depression angle of the image acquisition device 200, and h is the viewing angle range.

[0056] The image acquisition device 200 can also be configured with functions such as automatic focus and automatic exposure. The automatic focus is based on the principle of phase detection, and calculates the object distance by comparing different phase image information, and quickly adjusts the lens focal length; the automatic exposure is achieved by adjusting the exposure time and aperture size according to the image brightness distribution. Among them, the exposure of the image acquisition device 200 can be determined by the following formula:

[0057]

[0058] Among them, E is the exposure, I is the light intensity, t is the exposure time, and F is the aperture size.

[0059] The above-mentioned edge computing device 100 is a hardware or software system deployed at the network edge (close to the data source), used to complete computing, storage, and communication tasks locally or near the data source to reduce the dependence on remote cloud servers. For example, edge servers, edge gateways, edge routers, Internet of Things devices, etc., can be selected according to the actual situation by the edge computing device 100.

[0060] Among them, the edge computing device 100 is used to predict the corresponding behavior of the target image through a behavior prediction model. The behavior prediction model is obtained according to the behavior prediction model training method in the embodiments of the present application, and the behavior prediction model is stored in the edge computing device 100.

[0061] The edge computing device 100 is the core computing unit of the behavior prediction system. Among them, the edge computing device 100 can be built with a high-performance processor. Its operation speed is greater than or equal to 5.4 GHz.

[0062] In addition, an AI (Artificial Intelligence) acceleration chip can also be installed in the edge computing device 100 to accelerate the operation of AI algorithms and achieve fast processing of a large amount of image data.

[0063] In one embodiment, the edge computing device 100 processes data in a parallel computing and distributed computing manner. Among them, the calculation formula for the theoretical processing time of parallel computing can be: T1 is the processing time, N is the total number of tasks, and P is the degree of parallelism.

[0064] Distributed computing distributes tasks to multiple computing nodes to complete collaboratively. Through a reasonable task scheduling algorithm, such as dynamic priority scheduling based on the earliest deadline first strategy. In complex scenarios such as large factory monitoring, the task allocation can be dynamically adjusted according to the importance and data volume of different regions, giving priority to processing data in key production areas to ensure production safety.

[0065] Optionally, the edge computing device 100 can adopt a scheduling algorithm based on task priority.

[0066] For example, in the large factory monitoring scenario, a task urgency ranking can be set for each task. After monitoring that the production work in a certain area is completed, the task with the highest urgency can be scheduled to this area for production in a timely manner, thereby achieving the accuracy and flexibility of production task scheduling. In the above implementation process, the edge computing device 100 and the image acquisition device 200 are set in the behavior prediction system, and the edge computing device 100 and the image acquisition device 200 are connected. The target image obtained by the image acquisition device 200 can be directly transmitted to the edge computing device 100. Then, the edge computing device 100 predicts the behavior in the corresponding area of the target image according to the target image, reducing the data transmitted to the remote server, improving the efficiency of behavior monitoring, while reducing the network pressure and the security risk during data transmission.

[0067] In a possible implementation manner, the behavior prediction system further includes: a server.

[0068] Among them, the edge computing device 100 is connected to the server. The server can be a web server, a database server, a personal computer (PC), a tablet computer, a smart phone, a personal digital assistant (PDA), etc. The server can be selected according to the actual situation.

[0069] Optionally, the server is communicatively connected to one or more edge computing devices 100 via a network, and the edge computing device 100 is also used to transmit key feature data to the server. The data transmission between the server and the edge computing device 100 can use the TCP / IP protocol, and its transmission rate is greater than or equal to 1000 Mbps.

[0070] In one embodiment, the data is compressed and encrypted using the AES algorithm, and then the encrypted data is transmitted between the edge computing device 100 and the server.

[0071] When compressing the data, efficient video coding standards such as H.265 can be used, and the compression ratio can reach 1.5 - 2 times that of traditional H.264, which can greatly reduce the amount of data while ensuring the video quality.

[0072] When encrypting the data, the key length can be set to 128 bits, 192 bits, or 256 bits. Among them, the encryption process can be expressed as: C = E(K, P). C is the ciphertext, K is the key, and P is the plaintext.

[0073] To further ensure data security, quantum encryption technology can also be combined to further enhance data security. Quantum encryption is based on the principle of quantum key distribution. Utilizing the non-clonability of quantum states and the measurement collapse characteristics, it can ensure the absolute security of the key.

[0074] In addition, the load balancing technology based on Nginx can be adopted to evenly distribute a large number of requests to multiple server nodes, improving the system response speed and stability.

[0075] The key feature data is the main feature data used for behavior prediction. The key feature data can include profile data, pose data, location data, etc.

[0076] The server here is responsible for the unified management and scheduling of the system. The server is also used to receive the key feature data transmitted by the edge computing device 100, and conduct in-depth analysis and integration of the key feature data, and comprehensively evaluate and predict the overall situation of the monitored area based on big data analysis and artificial intelligence algorithms.

[0077] In one embodiment, the server has a remote control function and can adjust the parameters and allocate tasks for the image acquisition device 200 and the edge computing device 100. For example, when an emergency occurs in the monitored area, the camera shooting angle and focal length can be remotely adjusted through the server to obtain detailed information.

[0078] In the above implementation process, by setting up a server that is connected to the edge computing device 100, data that is difficult to process by the edge computing device 100 can be transmitted to the server for further processing, improving the data processing ability of the behavior prediction system. And when performing data transmission, only the key feature data is transmitted, which can reduce the amount of data transmitted, relieve the network bandwidth pressure, and improve the data transmission efficiency and stability.

[0079] In a possible implementation, the edge computing device 100 is configured to adjust the operating power consumption by adjusting the operating frequency and voltage of the processor.

[0080] Among them, the calculation formula for the operating power consumption is:

[0081] P = f × C × V 2 ;

[0082] Among them, P is the operating power consumption, f is the operating frequency of the processor, C is the load capacitance, and V is the voltage of the processor.

[0083] In the above implementation process, by setting the processor of the edge processing device to be able to adjust its operating power according to the operating frequency and voltage, the power consumption of the edge processing device can be reduced while improving the computing performance.

[0084] In a possible implementation, the edge computing device 100 includes a disk array.

[0085] Among them, the edge computing device 100 is configured to store the data for behavior prediction dispersedly on multiple independent disks in the disk array. The storage capacity of the disk array is greater than or equal to 512 GB.

[0086] It can be understood that when the network condition is poor or interrupted, the edge computing device 100 can store the data for behavior prediction locally in the edge computing device 100 and complete the behavior prediction independently. After the network returns to normal, the stored data can be transmitted to the server to ensure the continuity of the monitoring and prediction work.

[0087] In addition, the edge computing device 100 can have an automatic backup and recovery function. By using redundant storage technology (i.e., through an independent redundant disk array), the data is stored dispersedly on multiple disks to improve data reliability. When a certain disk fails, the data can be recovered through other disks.

[0088] In the above implementation process, by using a disk array for data storage, when a certain disk fails, the data can be recovered through other disks, which can prevent data loss while controlling costs.

[0089] The behavior prediction system in this embodiment can be used to execute each step in the various methods provided by the embodiments of the present application. The implementation process of the behavior prediction model training method will be described in detail through several embodiments below.

[0090] Please refer to Figure 2 , which is a flowchart of the behavior prediction model training method provided by the embodiments of the present application. The following will elaborate in detail on the Figure 2 specific process shown.

[0091] Step S201, training a basic model with sample data.

[0092] Among them, the basic model can be a general neural network model. For example, a convolutional neural network, a multi-layer perceptron, a recurrent neural network, etc. Of course, the basic model can also be an existing behavior prediction model. The model obtained by training the basic model includes a behavior prediction model. The basic model can be selected according to the actual situation. Here, the sample data refers to the data used for model training. For example, image data collected by an image acquisition device, behavior label data, etc. The sample data can be selected according to the actual situation.

[0093] The sample data can cover normal behavior patterns and abnormal behavior patterns to ensure data diversity. The sample data can be sourced from various behavior data (such as walking, running, fighting, abnormal staying, illegal vehicle driving, etc.) in multiple scenarios (such as security scenarios, intelligent traffic monitoring, commercial venue monitoring, etc.).

[0094] In one embodiment, before step S201, after obtaining the sample data, the sample data can be cleaned and labeled to remove noise data and mislabeled data in the sample data, and the behavior category of each sample can be clarified.

[0095] For ease of understanding, the following takes a convolutional neural network as the basic model as an example to show the training process of the basic model:

[0096] Exemplarily, the convolutional neural network can include a convolutional layer, a pooling layer, and a fully connected layer. Among them, the convolutional layer performs a convolution operation on the input sample data through a convolution kernel. The formula for the convolution operation is:

[0097] O(x,y) = ∑ i,j I(x + i,y + j)K(i,j);

[0098] Among them, the value range of (x,y) starts from (0,0) and ends at (H - h + 1,W - w + 1). H is the height of the output feature map, W is the width of the output feature map, h is the height of the convolution kernel, and K is the width of the convolution kernel. The value range of (i,j) is determined by the size of the convolution kernel.

[0099] For example, if the input sample data is an image with a resolution of 640×480 and the convolution kernel size is 3×3, the size of the output feature map is 638×478. Correspondingly, the range of values for (x, y) is from (0, 0) to (638, 478). And if the convolution kernel size is m×n, the range of values for i is from 0 to m - 1, and the range of values for j is from 0 to n - 1. Taking the common 3×3 convolution kernel as an example, the values of i are 0, 1, 2, and the values of j are also 0, 1, 2.

[0100] Among them, I is collected by an image acquisition device, which may be images of people, vehicles, etc. in a monitoring scenario. The convolution kernel is equivalent to a feature extractor, and its weight values determine the sensitivity to different features of the input sample data. Different convolution kernels can extract different features, such as edges, textures, etc.

[0101] Exemplarily, in a pedestrian behavior prediction scenario, a specific convolution kernel can capture key features such as the contour and action posture of a pedestrian. The convolution operation process is as follows: at each (x, y) position, the convolution kernel multiplies and sums the corresponding elements of the target area at (x, y) in the input image I to obtain the value of the output feature map at the (x, y) position, thereby realizing feature extraction.

[0102] For example, when detecting the abnormal staying behavior of a pedestrian, through the convolution operation, features with insignificant position changes of the pedestrian over a period of time can be extracted. The output feature map is the result of the convolution operation and contains the key features extracted from the input image. After being processed by multiple convolutional layers, the finally obtained feature map can be used as the input of a behavior classification model to determine which type of behavior the current behavior belongs to.

[0103] The target area can be the local area in the upper left corner, or the area around (x, y), or the starting position area, etc. The target area can be selected according to the actual situation.

[0104] The training process of the above basic model is only exemplary, and there may be certain differences in the model training methods corresponding to different types of basic models. The specific training process of this basic model can be adjusted according to the actual situation.

[0105] Step S202, use the trained model to predict the behavior of the target image and output the prediction result.

[0106] Among them, the target image is obtained by an image acquisition device in a set area.

[0107] Understandably, when training the base model, an iterative training method is usually adopted. That is, the base model is trained multiple times with multiple sample data. After each training of the base model, a target image can be input into the trained model, and based on the trained model, the behavior in the target image is predicted, and a prediction result is output.

[0108] Step S203, update the trained model according to the prediction result.

[0109] After the trained model outputs the prediction result, the corresponding adjustment parameters can be determined according to the multiple prediction results of the trained model, and the trained model is updated through the adjustment parameters to achieve the optimization of the model.

[0110] Step S204, continue to train the updated model with new sample data until the updated model training reaches the iteration condition.

[0111] After the trained model is updated, continue to input the new sample data that has not been used for model training into the updated model to continue training the updated model. After each training of the updated model, continue to predict the behavior of the target image through the trained model, output the prediction result, and update the trained model according to the prediction result. Iteratively train the model according to the above steps until the iteration condition is reached and the training stops.

[0112] The iteration condition here can be: reaching the iteration times, the prediction accuracy reaching the accuracy threshold, etc., and this iteration condition can be selected according to the actual situation.

[0113] In the above implementation process, by training the behavior prediction model on the edge computing device, and during the model training process, the model is iteratively updated through the prediction results of the trained model, the accuracy of the behavior prediction model in complex and changeable scenarios can be improved, and at the same time, data transmission can be reduced, and the processing efficiency and data security can be improved.

[0114] In a possible implementation manner, step S203 includes: determining a reward and punishment score according to the prediction result and the actual result; determining an optimization strategy through the reward and punishment score and the adjustment strategy; updating the prediction strategy of the trained model according to the optimization strategy.

[0115] Among them, the reward and punishment score is determined according to the accuracy of the prediction result, and this reward and punishment score can feedback the accuracy of the model behavior through setting a reward function.

[0116] Exemplarily, if the trained model accurately predicts an abnormal behavior and issues a warning in time, a positive reward can be given; if the prediction is wrong or the warning is not issued in time, a negative reward is given.

[0117] Furthermore, different scores can be given according to the impact degree in the corresponding scenario based on the prediction results of the trained model. For example, since a missed alarm may lead to the neglect of potential safety hazards and cause serious consequences, a relatively high negative reward score (e.g., -150) can be given; since a false alarm will waste the energy and time of security personnel and interfere with normal security work, but its harm is relatively smaller compared to a missed alarm, a relatively low negative reward score (e.g., -50) can be given; since a timely warning can enable staff to respond quickly and effectively avoid accidents, a medium positive reward (e.g., +100) can be given. Since accurately predicting the behavior at the next moment helps to make preparations in advance, a relatively low positive reward (e.g., +50) can be given. Since accurate long-term trend prediction helps to formulate more long-term strategies, a relatively high positive reward (e.g., +150) can be given.

[0118] The above settings of the reward and punishment scores are only exemplary, and the settings of the reward and punishment scores can be selected according to the actual situation.

[0119] By setting the correlation between the reward and punishment scores and the adjustment strategy, after determining the reward and punishment scores, the optimized strategy can be further determined through the reward and punishment scores, and then the prediction strategy of the trained model can be updated based on the optimized strategy.

[0120] Optionally, the correlation can be a control relationship between the reward and punishment scores and the optimized strategy, or an equation relationship constructed by the reward and punishment scores, the adjustment strategy, and the optimized strategy. The correlation can be selected according to the actual situation.

[0121] In one embodiment, the optimization objective of the optimized strategy is:

[0122] where γ is the discount factor and T2 is the time step.

[0123] For the convenience of understanding, the following shows the determination method of the reward and punishment parameters by listing several scenarios:

[0124] Security monitoring scenario:

[0125] In the security monitoring scenario, the main goal is to identify and warn of abnormal behaviors in a timely and accurate manner to ensure security.

[0126] Accurate warning reward: When the trained model accurately predicts abnormal behaviors (such as intrusion, fighting, etc.) and issues a warning in a timely manner, a positive reward can be given. For example, if an accurate warning is issued within a very short time (e.g., within 5 seconds) after an intrusion occurs, the reward and punishment score can be determined as: R = +100 points. Since a timely warning can enable security personnel to respond quickly and effectively avoid safety accidents, a relatively high reward can be given.

[0127] False negative penalty: If the trained model fails to detect abnormal behavior (i.e., a false negative occurs), a negative reward is given. For example, if an intrusion occurs but the trained model does not issue a warning, the reward and punishment score can be determined as: R = -200 points. Since false negatives may lead to the neglect of security risks and cause serious consequences, the penalty is relatively large.

[0128] False positive penalty: If the trained model issues a warning incorrectly (when there is actually no abnormal behavior), a certain negative reward can also be given (e.g., the reward and punishment score is: R = -50 points). Since false positives will waste the energy and time of security personnel and interfere with normal security work, but compared to false negatives, their harm is relatively small, so the penalty is also relatively light.

[0129] Intelligent transportation scenario:

[0130] In the intelligent transportation scenario, the focus is on ensuring smooth and safe traffic and reducing violations.

[0131] Correct recognition reward: When the trained model accurately recognizes normal traffic behaviors (such as a vehicle driving normally, a pedestrian crossing the road following traffic rules, etc.), a positive reward is given (e.g., the reward and punishment score is: R = +20 points). This helps the model strengthen its learning and recognition ability of normal behaviors.

[0132] Violation warning reward: If the trained model successfully predicts a traffic violation (such as running a red light, speeding, etc.) and issues a warning in a timely manner, a positive reward can be given (e.g., the reward and punishment score is: R = +80 points). Timely violation warnings can assist traffic management departments in taking measures to maintain traffic order.

[0133] False prediction penalty: If the trained model misjudges a normal behavior as a violation or fails to recognize a violation, a negative reward can be given. For example, the reward and punishment score for misjudgment is: R = -30 points, and the reward and punishment score for missed judgment is: R = -60 points. Since misjudgment will cause unnecessary interference to normal traffic participants and missed judgment will cause violations to go unhandled in a timely manner, different penalties can be given according to the degree of harm.

[0134] Commercial premises scenario:

[0135] In the commercial premises scenario, the goal is to analyze customer behavior and optimize business strategies.

[0136] Accurate behavior prediction reward: When the trained model accurately predicts a customer's purchase behavior, staying area, etc., a positive reward can be given. For example, accurately predicting that a customer will purchase a certain product, which helps the merchant prepare inventory management and marketing in advance, the reward and punishment score can be: R = +50 points.

[0137] Error prediction penalty: If the trained model has a large deviation in predicting customer behavior, a negative reward can be given. For example, if the purchase intention of a customer is mispredicted, resulting in overstocking of goods, the reward and punishment score can be: R = -40 points.

[0138] Long-term behavior trend prediction reward: If the trained model can accurately predict the long-term behavior trends of customers (such as changes in customer consumption frequency, preference transfer, etc.), and accurate long-term trend prediction helps merchants formulate more long-term business strategies, a relatively high positive reward can be given (for example, R = +150 points).

[0139] The determination of the above reward and punishment scores is only exemplary, and the determination of the reward and punishment scores can be selected according to the actual situation.

[0140] The prediction strategy here refers to the methods, techniques, steps, etc. adopted by the behavior prediction model during the behavior prediction process. The optimization strategy refers to the prediction strategy that can improve the accuracy and stability of the behavior prediction model.

[0141] It can be understood that when the behavior prediction model performs behavior prediction on the target image, each step in the prediction process can correspond to one or more adjustment strategies, and these processing strategies can be set in advance. To improve the accuracy of the behavior prediction model, it is necessary to determine the adjustment strategy corresponding to the highest prediction result accuracy of the behavior prediction model from these adjustment strategies. By determining the corresponding reward and punishment scores for the prediction results in each iteration process, and then the adjustment strategy associated with it can be determined according to the reward and punishment scores to determine the corresponding optimization strategy, and then the prediction strategy of the trained model can be updated through the optimization strategy, so that the prediction strategy of the behavior prediction model obtained after the training is the optimal strategy.

[0142] In one embodiment, a neural network architecture can be used as the policy network. For example, a multi-layer perceptron (MLP) or a convolutional neural network (CNN) can be used, with the input being the image features collected by the image acquisition device and the output being the predicted behavior category or action.

[0143] In the behavior prediction system, if CNN is used for feature extraction, MLP can be connected subsequently as the policy network to output the predicted behavior. Among them, during the initialization of parameters: randomly initialize the weights and biases of the policy network, and set the initial exploration rate (such as the ε value in the ε-greedy policy) to balance the exploration and exploitation capabilities of the model. In the environment interaction: let the model run in the monitoring scenario, make behavior decisions (such as predicting the behavior category) according to the output of the current policy network, and observe the reward and punishment scores feedback by the environment.

[0144] Then, calculate the policy gradient. During the interaction between the model and the environment, record the state at each step (such as the current image features), the actions taken (predicted behaviors), and the obtained reward and punishment scores to form trajectory data.

[0145] Furthermore, according to the goal of reinforcement learning, calculate the cumulative discounted reward from the current time step to the future. where γ is the discount factor, R t+k is the reward at future time steps, T is the total number of time steps, and t is the current time step. Using the policy gradient theorem, calculate the gradient of the policy network parameter θ with respect to the cumulative reward. Taking an algorithm based on policy gradient (e.g., REINFORCE algorithm) as an example, its policy gradient formula is where π θ (a t |s t ) is the probability of taking action a t in state s t .

[0146] Next, update and optimize the policy; select and use stochastic gradient descent (SGD) or its variants (such as Adam, RMSProp, etc.) as the optimization algorithm through the optimization algorithm, and update the parameters of the optimized policy according to the calculated policy gradient. Use the parameter update to update the weights and biases of the optimized policy according to the rules of the optimization algorithm. Taking the Adam algorithm as an example, its update formula is where ω t is the current parameter, α is the learning rate, m t is the first moment estimate of the gradient, v t is the second moment estimate of the gradient, and ε is a small constant to prevent the denominator from being zero. After updating and optimizing the policy, adjust the behavior of the model according to the setting of the exploration rate. If the ε-greedy strategy is adopted, randomly select actions with probability ε for exploration, select actions according to the output of the optimized policy with probability 1 - ε for exploitation, and gradually reduce the value of ε so that the model relies more on the learned policy in the later stage of training.

[0147] Final repeated training and policy adjustment: Repeat the above process of environment interaction, calculation of policy gradients, and update and optimization of policies through multiple rounds of training. In each round of training, the model further optimizes the optimized policy based on the newly obtained reward and punishment scores. For example, in the security monitoring scenario, when the model accurately predicts abnormal behaviors (such as intrusion, fighting, etc.) during training and issues a warning in a timely manner, a positive reward is given. For instance, if a warning is accurately issued within a very short time (e.g., within 5 seconds) after an intrusion occurs, the reward R = +100 points. This is because a timely warning enables security personnel to respond quickly and effectively avoid the occurrence of safety accidents, so a relatively high reward is given. If the model fails to detect an abnormal behavior, i.e., a false negative occurs, a negative reward is given. For example, if an intrusion occurs but the model does not issue a warning, the reward R = -200 points. False negatives may lead to the neglect of potential safety hazards and cause serious consequences, so the punishment is relatively severe. If the model issues a warning incorrectly (when there is actually no abnormal behavior), a certain negative reward is also given, such as R = -50 points. False positives waste the energy and time of security personnel and interfere with normal security work, but compared to false negatives, their harm is relatively minor, so the punishment is also relatively light.

[0148] Optionally, after training the behavior prediction model, the performance of the behavior prediction model can also be evaluated regularly using a validation set or a test set, and the changes in indicators such as accuracy and recall rate can be observed. If the performance improvement of the behavior prediction model is not obvious or overfitting occurs, hyperparameters such as the learning rate and discount factor are adjusted, or the structure of the optimized policy is adjusted, such as increasing the number of network layers, adjusting the number of neurons, etc., to optimize the model's response to reward and punishment scores and the effect of policy adjustment.

[0149] In the above implementation process, the prediction results of the trained model are scored based on the reward and punishment scores, and then the corresponding optimized policy is determined according to the reward and punishment scores, and the prediction policy of the trained model is updated, which can improve the accuracy of the prediction policy in the model, and further improve the accuracy of the behavior prediction model.

[0150] In a possible implementation manner, step S203 includes: determining the evaluation indicators of the trained model according to the multiple prediction results of the behavior prediction of the target image by the trained model; adjusting the hyperparameters in the trained model according to the evaluation indicators and the hyperparameter tuning algorithm.

[0151] Among them, the evaluation indicators may include: accuracy, recall rate, comprehensive value, etc. The accuracy is used to reflect the proportion of correct predictions by the model, the recall rate is used to measure the model's ability to capture the entire sample, and the comprehensive value takes into account the values of the accuracy and recall rate.

[0152] In one embodiment, during the process of behavior prediction, a confusion matrix can be constructed first. Taking binary classification (such as normal behavior and abnormal behavior classification) as an example, there will be four combinations between the model prediction result and the actual situation: True Positive (TP), False Positive (FP), True Negative (TN), and False Negative (FN). In an actual scenario, if the model correctly predicts an abnormal behavior once, this is a true positive; if the model mispredicts a normal behavior as an abnormal behavior, this belongs to a false positive; correctly identifying a normal behavior is a true negative; and misjudging an abnormal behavior as a normal behavior is a false negative. For multi-classification problems (such as identifying multiple behaviors like "walking", "running", "fighting", etc.), the confusion matrix will be extended to a \(C\times C\) matrix, where \(C\) is the number of classes, and each element in the matrix represents the number of samples that are actually in one class but are predicted to be in another class.

[0153] Secondly, use model evaluation metrics such as accuracy, recall rate, and F1 value to strictly evaluate the trained model. Among them, accuracy represents the proportion of the number of samples correctly predicted by the model to the total number of samples, and the formula is: where Accuracy is the accuracy rate, TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative. Suppose in a test set containing 100 behavior samples, the model correctly predicts 85 samples (among them, 30 true positives and 55 true negatives), and incorrectly predicts 15 samples (5 false positives and 10 false negatives), then the accuracy rate is That is, 85%. In practical applications, the accuracy rate can intuitively reflect the overall prediction correctness of the model.

[0154] The calculation of the recall rate (completeness rate) refers to the proportion of true positives among all actual positive examples, and the formula is: where Recall is the recall rate, TP is the true positive, and TN is the true negative. Continuing with the above example, if there are 40 actual abnormal behavior samples (i.e., TP + FN = 40), and the model correctly predicts 30 of them (TP = 30), then the recall rate is That is, 75%. The recall rate is very crucial in measuring the ability of the model to capture positive samples (such as abnormal behaviors). In scenarios such as security monitoring, a higher recall rate means that more potential dangerous behaviors can be detected as much as possible to avoid missed reports.

[0155] The comprehensive value is an index that comprehensively considers the accuracy rate and the recall rate, and its calculation formula is: Among them, F1 is the comprehensive value, Accuracy is the accuracy, and Recall is the recall. Substituting the previously calculated accuracy of 0.85 and recall of 0.75 into the formula, we can get The F1 value can more comprehensively evaluate the performance of the model. When both the accuracy and recall are high, the F1 value will be high, which helps to make more reasonable comparisons and selections between different models.

[0156] For multi-classification problems, macro-average and micro-average can be used for calculation. Macro-average is to calculate the accuracy, recall and comprehensive value for each category separately, and then find the average value; micro-average is to summarize the true positives, false positives, false negatives, etc. of all categories, and then calculate the corresponding indicators according to the binary classification formula. Taking the macro-average comprehensive value as an example, assuming that the model needs to identify three types of behaviors, A, B, and C, and the comprehensive values ​​of A, B, and C are calculated to be 0.8, 0.7, and 0.6, respectively, then the macro-average comprehensive value is

[0157] The calculation of the above parameters such as accuracy, recall, and comprehensive value is only exemplary, and the calculation of evaluation indicators can be selected according to actual conditions.

[0158] After the behavior of the target image is predicted by the trained model, the prediction results of each iteration of the model can be stored. When determining the evaluation index, it can be determined based on the stored prediction results.

[0159] Optionally, the hyperparameter tuning algorithm may be a random search, a grid search, etc. The hyperparameter tuning algorithm may be selected according to actual conditions.

[0160] Random search is a hyperparameter tuning method that searches for the optimal hyperparameter combination by randomly sampling in a given hyperparameter space.

[0161] Grid search is a hyperparameter tuning method that searches for the best performance of a model by traversing all predefined parameter candidate values. Its core logic is to build a multidimensional grid in the parameter space, evaluate the effect of each parameter combination one by one, and finally select the combination with the best performance on the validation set.

[0162] It should be understood that when the hyperparameter tuning algorithm determines the hyperparameters, the performance of each set of hyperparameters can be evaluated according to the value of the evaluation index to determine the optimal hyperparameters.

[0163] The following uses a convolutional neural network (CNN) for behavior prediction combined with random search and grid search hyperparameter tuning as an example to show in detail the specific implementation of hyperparameter adjustment in the embodiments of the present application:

[0164] First, determine the hyperparameter range. For the CNN model, it is necessary to determine the value range of hyperparameters. Among them, the convolutional kernel size can be set to [3, 5, 7]. These sizes are commonly used in image feature extraction. Small-sized convolutional kernels are good at capturing local details, while large-sized convolutional kernels can obtain more extensive context information. The stride is set to [1, 2]. A stride of 1 can better preserve image details, and a stride of 2 can accelerate the downsampling speed and reduce the computational cost. The learning rate is selected from [0.001, 0.01, 0.1]. Different learning rates determine the step size of parameter updates during model training. An appropriate learning rate can make the model converge faster. The number of hidden layer nodes is set to [64, 128, 256]. The number of hidden layer nodes affects the learning ability of the model. More nodes can learn more complex features but may also lead to overfitting.

[0165] Secondly, select accuracy, recall, and the comprehensive value as evaluation metrics. These metrics can evaluate the model performance from different perspectives. Accuracy reflects the proportion of correct predictions of the model. Recall measures the ability of the model to capture positive samples. The comprehensive value comprehensively considers accuracy and recall and more comprehensively evaluates the performance of the model in the behavior prediction task. Then, set the number of searches. Suppose the number of searches is set to 50 times. Too few search times may not be able to fully explore the hyperparameter space and it is difficult to find a better solution. Too many search times will increase the computational cost and time.

[0166] The random search process is to randomly select a set of hyperparameters from the hyperparameter range each time. For example, in the first random selection, the convolutional kernel size is 5, the stride is 1, the learning rate is 0.01, and the number of hidden layer nodes is 128. Train the model with this set of hyperparameters and calculate the evaluation metrics on the validation set. Repeat this process 50 times and record the hyperparameter combinations and corresponding evaluation metrics for each training.

[0167] Finally, determine the optimal hyperparameters: Compare the evaluation metrics obtained from 50 searches and select the hyperparameter combination that maximizes the comprehensive value as the optimal hyperparameters. Suppose during the search process, when the hyperparameter combination of a certain training is that the convolutional kernel size is 3, the stride is 2, the learning rate is 0.001, and the number of hidden layer nodes is 256, the model obtains the highest comprehensive value of 0.85 on the validation set. Then this set of hyperparameters is the optimal hyperparameters obtained by random search.

[0168] For grid search of hyperparameters, the range of hyperparameters needs to be determined first. Similar to random search, determine the range of hyperparameters for the CNN model, such as the convolutional kernel size [3, 5, 7], stride [1, 2], learning rate [0.001, 0.01, 0.1], and the number of nodes in the hidden layer [64, 128, 256]. Also, select accuracy, recall, and the comprehensive value as evaluation metrics. Permute the values of each hyperparameter to generate all possible hyperparameter combinations. In this example, the number of hyperparameter combinations is 3×2×3×3 = 54. Train the model with each hyperparameter combination in turn and calculate the evaluation metrics on the validation set. For example, first train the model with the hyperparameter combination of convolutional kernel size 3, stride 1, learning rate 0.001, and the number of nodes in the hidden layer 64, and record its accuracy, recall, and comprehensive value on the validation set; then continue to train and evaluate with the next set of hyperparameters until all 54 combinations have been trained and evaluated. Finally, compare the evaluation metrics corresponding to all hyperparameter combinations and select the hyperparameter combination that maximizes the comprehensive value. Suppose after grid search, it is found that when the convolutional kernel size is 5, stride is 1, learning rate is 0.01, and the number of nodes in the hidden layer is 256, the comprehensive value of the model on the validation set is the highest, reaching 0.88. Then this set of hyperparameters is the optimal hyperparameters obtained by grid search. In actual operation, by continuously trying different hyperparameter combinations and selecting the optimal parameter settings according to the evaluation metrics, the overall performance of the model can be improved.

[0169] In the above implementation process, by adjusting the hyperparameters in the trained model according to the evaluation metrics and hyperparameter tuning algorithms, the optimal parameter settings corresponding to the model can be determined, and the overall performance of the model can be improved.

[0170] In a possible implementation manner, step S203 includes: obtaining the first model parameters of the trained related model; adjusting the second model parameters of the trained model through the first model parameters.

[0171] The related model here refers to a model related to the behavior prediction model. For example, a general large model, an existing behavior prediction model, a model in a similar field, etc. The related model can be selected according to the actual situation.

[0172] Exemplarily, for behavior prediction in a security scenario, since it involves the recognition and prediction of human behaviors in surveillance footage and is related to image classification and action recognition, a convolutional neural network (CNN) model pre-trained on a large image classification dataset (such as ImageNet), such as ResNet50, can be selected. This model performs well in image classification tasks, has strong feature extraction capabilities, can learn rich image features, and provides strong support for subsequent behavior prediction.

[0173] Understandably, for a model in the initial stage of training, since the model parameters are usually determined based on experience or randomly, it is possible that these model parameters have little relevance to behavior prediction. In this case, the training of the model is difficult and time-consuming.

[0174] If the first model parameters in the existing related model are directly migrated to the trained model, the trained model can be fine-tuned using a small amount of target task data, which can accelerate the convergence speed of the model, improve the training efficiency, and reduce the dependence on a large amount of labeled data.

[0175] For example, when training a behavior prediction model for security monitoring, a model pre-trained on an image classification dataset can be migrated, and then the second model parameters in the model can be fine-tuned in combination with a small amount of data in the security scenario.

[0176] Among them, the migration method can include overall migration and partial migration, and the first model parameters can be determined according to the migration method.

[0177] Overall migration can directly migrate all the parameters of the related model to the trained model, freeze the parameters of the related model, and only train the newly added classification layer or specific task layer. This method is applicable to the situation where the source task and the target task are highly similar and the amount of target task data is small. For example, in the image classification task in certain specific fields, if it is close to the training field of the related model and the amount of data is limited, overall migration can be used to quickly build a model and obtain good results.

[0178] Partial migration is to select some layers of the related model for migration, usually the underlying feature extraction layers, because the features learned by these layers are general image features, such as edges, textures, etc. For the high-level classification layer, it is reconstructed according to the number and characteristics of the categories of the target task. This method has high flexibility and is applicable to the situation where there are certain differences between the source task and the target task, but there are also some commonalities.

[0179] It should be understood that for overall migration, the weights of the related model can be loaded first, and the parameters of the layers with the same name and consistent dimensions can be directly assigned. For partial migration, only the parameters of the selected layers can be migrated, and then the parameters are matched and migrated to ensure that the parameters of the related model can be correctly matched with the layer structure of the trained model.

[0180] Based on the selected relevant model, the model structure can be adjusted according to the requirements of the target task. Generally, one or more fully connected layers are added at the end of the relevant model as the classification layer. If the target task is a multi-classification problem, the output dimension of the classification layer should match the number of classes; if it is a regression problem, one or more continuous values are output. Taking the behavior prediction task as an example, assuming that 10 different behaviors such as "walking", "running", "fighting" etc. need to be recognized, after migrating some layers of the ResNet50 model, a fully connected layer is added, its output dimension is set to 10, and the softmax function is used as the activation function to convert the model output into probabilities of each class, so as to achieve the classification of different behaviors.

[0181] During the fine-tuning process, training parameters such as the learning rate can be adjusted according to the actual situation. Generally speaking, the learning rate should be smaller than that when training from scratch to avoid overfitting. At the same time, some training techniques can be adopted, such as early stopping method, data augmentation, etc. In the behavior prediction task in the embodiment of the present application, the collected sample data in the monitoring scenario is used to fine-tune the trained model. During the training process, indicators such as the accuracy rate and recall rate of the model on the validation set are observed, and the training is stopped when the indicators no longer improve to prevent overfitting.

[0182] The model after parameter migration and fine-tuning consists of the migrated relevant model part and the specific layers added for the target task. Taking the behavior prediction model migrated based on ResNet50 as an example, the model structure from input to output is as follows: some convolutional layers and pooling layers of ResNet50 (used to extract general image features), followed by a global average pooling layer (to convert the feature map into a feature vector with a fixed length), then a newly added fully connected layer (used to map the feature vector to the behavior category space), and finally a softmax layer (to convert the output of the fully connected layer into probabilities of each class to achieve behavior classification). This structure makes full use of the feature extraction ability of the relevant model and combines the classification layer designed for the target task, and can effectively complete the behavior prediction task.

[0183] In the above implementation process, by using the first model parameters of the trained relevant model to adjust the second model parameters of the trained model, the fine-tuning of the model can be realized with a small amount of data, which can improve the model convergence speed, improve the training efficiency, and at the same time reduce the dependence on data annotation.

[0184] Please refer to Figure 3 , which is the flowchart of the behavior prediction method provided by the embodiment of the present application. The following will elaborate on the Figure 3 specific process shown in detail.

[0185] Step S301, input the target image into the behavior prediction model.

[0186] Step S302: Predict the behavior corresponding to the target image through the behavior prediction model.

[0187] Among them, the behavior prediction model is obtained according to the behavior prediction model training method in the above embodiments.

[0188] The target image here can be obtained in real time and input into the behavior prediction model in real time, and then the behavior corresponding to the target image can be predicted in real time through the behavior prediction model.

[0189] The behavior prediction model is stored in the edge computing device, and the behavior prediction is performed through the edge computing device.

[0190] In one embodiment, after step S302, the method further includes: issuing an abnormal warning when the behavior corresponding to the target image is an abnormal behavior.

[0191] In the above implementation process, the behavior prediction of the target image is performed through the behavior prediction model stored in the edge computing device, which can reduce a large amount of data transmission, reduce the requirement for network bandwidth, effectively avoid network congestion, and improve the stability of behavior prediction in a complex network environment. Even in areas with poor network conditions, real-time monitoring and behavior prediction can be achieved, and the accuracy of behavior prediction can be improved.

[0192] In one possible implementation manner, step S302 includes: determining the historical data and current behavior data corresponding to the target image; inputting the historical data and current behavior data into an autoregressive model to generate a prediction value; classifying through the softmax function and the prediction value to determine the probability value of each behavior; determining the behavior with the largest probability value as the behavior corresponding to the target image.

[0193] Among them, the historical data refers to the target image data obtained before the current moment. The current behavior data refers to the behavior data in the target image at the current moment.

[0194] In one embodiment, the mathematical expression of the prediction value can be:

[0195]

[0196] Among them, p is the order of the autoregressive part, q is the order of the moving average part, is the autoregressive coefficient (i = 1, 2..., p), θ j is the moving average coefficient (j = 1, 2..., q), ε t is a white noise sequence, Y t is the behavior feature value at the current moment.

[0197] The above autoregressive coefficients need to satisfy certain conditions. For example, for a stationary autoregressive process, the roots of its characteristic equation must all be outside the unit circle. The roots of the characteristic equation of the moving average coefficients should be outside the unit circle.

[0198] Both p and q are non-negative integers, and the values of p and q are determined according to the characteristics of the data and actual needs. Different combinations of p and q can be tried, and the information criterion can be used to select the optimal model order.

[0199] The white noise sequence is usually assumed to follow a normal distribution with a mean of 0 and a variance of σ 2 That is, ε t ~N(0,σ 2 ).

[0200] In the above mathematical expression of the predicted value, the autoregressive part reflects the linear relationship between the current behavior characteristics and the behavior characteristics at the previous p time instants, and reflects the self-evolution law and trend of the behavior characteristics. For example, if analyzing the movement trajectory of personnel in a monitoring scenario, the autoregressive part can capture the influence of the previous position and speed of the personnel on the current position. The moving average part represents that the current behavior characteristics are affected by the random disturbances (white noise) at the previous q time instants and the comprehensive effect of the random disturbance at the current time instant. In actual behavior prediction, these random disturbances can be understood as the influence of some unpredictable sudden factors on the behavior, and the moving average part incorporates these factors into the model, enabling the model to better fit the actual data.

[0201] By analyzing the historical behavior characteristic data, the time series law of the behavior is mined, so as to predict the future behavior trend. For example, when analyzing the personnel flow behavior in a shopping mall, according to the staying positions and movement trajectories of the personnel in the past period of time, predict their next action direction.

[0202] It should be understood that after analyzing the data in the target image and determining the corresponding behavior type, the behavior type can also be classified to determine the probability value corresponding to each behavior type, and then determine that the behavior with the highest probability value is the behavior corresponding to the target image.

[0203] Among them, classification can be performed through the softmax function, and its formula is:

[0204]

[0205] Among them, x is the model output, y is the category, and k is the total number of categories.

[0206] Here is always greater than 0. The sum of K numbers greater than 0 is also greater than 0. Therefore, the value range of P(y = k|x) is (0, 1). At the same time, because that is, the sum of all class probabilities is 1.

[0207] The softmax function can transform the raw values (usually unnormalized scores or log odds) output by the behavior prediction model into probability values corresponding to each category. These probability values reflect the relative likelihood of a sample belonging to each category given the input. For example, in a real-time behavior prediction image acquisition device, after the behavior prediction model analyzes the behavior in the surveillance video and obtains a set of raw outputs, through the processing of the softmax function, the probabilities of the current behavior belonging to different categories such as "walking", "running", "fighting", etc. can be obtained. The larger the value, the higher the likelihood that the behavior belongs to the corresponding category.

[0208] In the above implementation process, when performing behavior prediction, determining the behavior with the largest probability value as the behavior corresponding to the target image can improve the accuracy of behavior prediction.

[0209] Based on the same inventive concept, an embodiment of the present application also provides a behavior prediction model training device corresponding to the behavior prediction model training method. Since the principle of solving problems by the device in the embodiment of the present application is similar to that of the foregoing behavior prediction model training method embodiment, the implementation of the device in this embodiment can refer to the description in the embodiment of the above method, and the repeated parts will not be elaborated.

[0210] Please refer to Figure 4 , which is a schematic diagram of the functional modules of the behavior prediction model training device provided by the embodiment of the present application. Each module in the behavior prediction model training device in this embodiment is used to execute each step in the above method embodiment. The behavior prediction model training device includes a training module 401, a first prediction module 402, and an update module 403; among them,

[0211] The training module 401 is used to train the basic model through sample data; among them, the model obtained by training the basic model includes a behavior prediction model.

[0212] The first prediction module 402 is used to predict the behavior of the target image through the trained model and output a prediction result; among them, the target image is obtained by an image acquisition device in a set area.

[0213] The update module 403 is used to update the trained model according to the prediction result.

[0214] The training module 401 is further used to continue training the updated model through new sample data until the updated model training reaches the iteration condition.

[0215] In a possible implementation, the updating module 403 is further configured to: determine a reward and punishment score according to the prediction result and the actual result; wherein, the reward and punishment score is determined according to the accuracy of the prediction result; determine an optimization strategy through the reward and punishment score and an adjustment strategy; and update the prediction strategy of the trained model according to the optimization strategy.

[0216] In a possible implementation, the updating module 403 is further configured to: determine an evaluation index of the trained model according to multiple prediction results of the trained model for behavior prediction of the target image; and adjust hyperparameters in the trained model according to the evaluation index and a hyperparameter tuning algorithm.

[0217] In a possible implementation, the updating module 403 is further configured to: obtain first model parameters of a trained related model; and adjust second model parameters of the trained model through the first model parameters.

[0218] Based on the same application concept, an embodiment of the present application further provides a behavior prediction model training device corresponding to the behavior prediction model training method. Since the principle of the device in the embodiment of the present application for solving problems is similar to that of the foregoing embodiment of the behavior prediction model training method, the implementation of the device in this embodiment can refer to the description in the embodiment of the above method, and repeated parts will not be described again.

[0219] Please refer to Figure 5 , which is a schematic diagram of functional modules of the behavior prediction device provided by the embodiment of the present application. Each module in the behavior prediction device in this embodiment is used to execute each step in the above method embodiment. The behavior prediction device includes an input module 501 and a second prediction module 502; wherein,

[0220] The input module 501 is configured to input the target image into a behavior prediction model; wherein, the behavior prediction model is obtained according to the behavior prediction model training method in the above embodiment.

[0221] The second prediction module 502 is configured to predict the behavior corresponding to the target image through the behavior prediction model.

[0222] In a possible implementation, the second prediction module 502 is further configured to determine historical data and current behavior data corresponding to the target image; input the historical data and the current behavior data into an autoregressive model to generate a predicted value; classify through a softmax function and the predicted value to determine the probability value of each behavior; and determine the behavior with the largest probability value as the behavior corresponding to the target image.

[0223] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the behavior prediction model training method and / or the behavior prediction method described in the above method embodiments.

[0224] A computer program product of the prediction model training method and / or the behavior prediction method provided by the embodiments of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the prediction model training method and / or the behavior prediction method described in the above method embodiments. For details, refer to the above method embodiments and will not be elaborated here.

[0225] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0226] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0227] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.

[0228] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0229] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or replacements, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A method for training a behavior prediction model, characterized in that, Applied to an edge computing device, the method includes: Training a basic model with sample data; wherein, the model obtained by training the basic model includes a behavior prediction model; Predicting the behavior of a target image through the trained model and outputting a prediction result; wherein, the target image is acquired by an image acquisition device in a set area; Updating the trained model according to the prediction result; Continuing to train the updated model with new sample data until the training of the updated model reaches the iteration condition.

2. The method according to claim 1, wherein The updating the trained model according to the prediction result includes: Determining a reward and punishment score according to the prediction result and the actual result; wherein, the reward and punishment score is determined according to the accuracy of the prediction result; Determining an optimization strategy through the reward and punishment score and an adjustment strategy; Updating the prediction strategy of the trained model according to the optimization strategy.

3. The method according to claim 1 or 2, characterized in that, The updating the trained model according to the prediction result includes: Determining an evaluation index of the trained model according to multiple prediction results of the trained model for behavior prediction of the target image; Adjusting hyperparameters in the trained model according to the evaluation index and a hyperparameter tuning algorithm.

4. The method according to claim 1 or 2, characterized in that, The updating the trained model according to the prediction result includes: Obtaining first model parameters of a trained related model; Adjusting second model parameters of the trained model through the first model parameters.

5. A behavior prediction method, characterized in that, Applied to an edge computing device, the method includes: Inputting the target image into a behavior prediction model; wherein, the behavior prediction model is obtained according to the behavior prediction model training method of any one of claims 1-4; Predicting the behavior corresponding to the target image through the behavior prediction model.

6. The method according to claim 5, wherein The predicting the behavior corresponding to the target image through the behavior prediction model includes: Determining historical data and current behavior data corresponding to the target image; Inputting the historical data and the current behavior data into an autoregressive model to generate a prediction value; Classifying through a softmax function and the prediction value to determine the probability value of each behavior; Determining the behavior with the largest probability value as the behavior corresponding to the target image.

7. A behavior prediction system, characterized in that, Includes: An image acquisition device and an edge computing device; The edge computing device is connected to the image acquisition device; The image acquisition device is configured to acquire a target image and transmit the target image to the edge computing device; The edge computing device is configured to predict the behavior corresponding to the target image through a behavior prediction model; wherein, the behavior prediction model is obtained according to the behavior prediction model training method of any one of claims 1-4, and the behavior prediction model is stored in the edge computing device.

8. The system according to claim 7, wherein Further includes: A server; The edge computing device is connected to the server; The edge computing device is further configured to transmit key feature data to the server; Wherein, the key feature data is main feature data for behavior prediction; the key feature data includes contour data, pose data, and position data.

9. The system according to claim 7, wherein The edge computing device is configured to adjust the operating power consumption by adjusting the working frequency and voltage of the processor; Wherein, the calculation formula of the operating power consumption is: P = f × C × V 2 ; Among them, P is the operating power consumption, f is the operating frequency of the processor, C is the load capacitance, and V is the voltage of the processor.

10. The system according to claim 7, wherein The edge computing device includes a disk array; The edge computing device is configured to dispersedly store the data for behavior prediction on a plurality of independent disks in the disk array.

Citation Information

Cited By

  • AI glasses data high-speed read-write method and system based on storage chip

    CN121116204A