Target Behavior Recognition Method, Target Behavior Recognition Device and Computer Storage Medium

By combining the multimodal information fusion method of node information sequence and reference monitoring video frame, the accuracy of monkey car equipment behavior recognition is solved, and the precise control of monkey car and power saving is achieved to meet passenger needs.

CN115205728BActive Publication Date: 2025-07-22ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210619010.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-07-22
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

When existing video surveillance systems identify the behavior of monkey car equipment and its operating scenarios, the single-modal information recognition accuracy is low and are easily affected by the environment, resulting in waste of electricity and increased operating costs.

Method used

The multimodal information fusion method is adopted, combining the node information sequence and the reference monitoring video frame, and the space-time graph convolution network and the improved YOLO V3 network are used to identify the target behavior, and determine whether the target object has preset behavior through the weighted sum of the node information sequence and the reference monitoring video frame.

Benefits of technology

It improves the accuracy and speed of target behavior recognition, ensures the accurate operation of monkey car, saves power resources, and meets passenger needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205728B_ABST
    Figure CN115205728B_ABST
Patent Text Reader

Abstract

The present application discloses a target behavior recognition method, a target behavior recognition device, and a computer storage medium. The target behavior recognition method includes: obtaining a tracking video of a target object, where the tracking video includes multiple monitoring video frames containing the target object; determining a sequence of joint point information of the target object based on the multiple monitoring video frames, where the sequence of joint point information includes the joint point information corresponding to each monitoring video frame in the multiple monitoring video frames; and determining a reference monitoring video frame of the target object based on the multiple monitoring video frames; using the sequence of joint point information and the reference monitoring video frame to determine whether the target object has a preset behavior. The target behavior recognition method of the present application can combine two types of modal information, namely the sequence of joint point information and the reference monitoring video frame, to jointly recognize the behavior of the target in the monitoring video, and improve the recognition accuracy of the target behavior recognition method and accelerate the recognition speed by using the combined modal information method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a target behavior recognition method, a target behavior recognition device, and a computer storage medium. Background Art

[0002] The full name of the Monkey Car is the Mine Aerial Passenger Ropeway, which is mainly used for assisting the transportation of workers in underground mines. Currently, in most mines, the Monkey Car equipment runs 24 hours a day or at fixed time periods. If it runs 24 hours a day, a lot of electricity will be wasted, increasing operating costs; if it runs at fixed time periods, some workers who cannot get on the Monkey Car in time can only walk in and out.

[0003] Therefore, the video monitoring system needs to monitor the monkey car equipment and its operation scene in real time, so as to automatically control the operation of the monkey car according to the monitoring situation. However, the current video monitoring system only recognizes the behavior of the monkey car and the staff through single-mode information such as image information, which is easily affected by the environment, resulting in low recognition accuracy. Summary of the invention

[0004] The present application provides a target behavior recognition method, a target behavior recognition device and a computer storage medium.

[0005] A technical solution adopted by the present application is to provide a target behavior recognition method, the target behavior recognition method comprising:

[0006] Acquire a tracking video of a target object, wherein the tracking video includes a plurality of monitoring video frames including the target object;

[0007] Based on the multiple surveillance video frames, determining a joint point information sequence of the target object, the joint point information sequence including joint point information corresponding to each surveillance video frame in the multiple surveillance video frames; and

[0008] Based on the multiple surveillance video frames, determining a reference surveillance video frame of the target object, wherein the reference surveillance video frame is obtained based on at least one surveillance video frame of the multiple surveillance video frames;

[0009] The joint point information sequence and the reference monitoring video frame are used to determine whether the target object has a preset behavior.

[0010] Wherein, the determining whether the target object has a preset behavior by using the joint point information sequence and the reference monitoring video frame includes:

[0011] Determine the sequence action information of the target object by using the joint point information sequence; and

[0012] Determine the behavior reference information of the target object by using the reference monitoring video frame;

[0013] Based on the sequence action information and the behavior reference information, determine whether the target object has the preset behavior.

[0014] Among them, the determining the sequence action information of the target object by using the sequence of joint point information includes:

[0015] Determine the preset sequence action that matches the sequence of joint point information among multiple preset sequence actions as the sequence action information of the target object.

[0016] Among them, the multi-frame monitoring video frames are collected for a transportation vehicle; the sequence action information is determined based on at least one of the position distribution information of the target object and the transportation vehicle, and the posture of the target object.

[0017] Among them, the transportation vehicle includes a vehicle with a steel cable as a track, and the multi-frame monitoring video frames are collected for the area where the target object enters or leaves the transportation vehicle.

[0018] Among them, the determining the preset sequence action that matches the sequence of joint point information among multiple preset sequence actions as the sequence action information of the target object includes:

[0019] Based on the sequence of joint point information, obtain the target sequence action feature of the target object;

[0020] Use a spatio-temporal graph convolutional network to update the target sequence action feature;

[0021] Compare the updated target sequence action feature with the action features of multiple preset sequence actions respectively, and determine the preset sequence action with the highest feature similarity as the sequence action information of the target object.

[0022] Among them, the determining the reference monitoring video frame of the target object based on the multi-frame monitoring video frames includes:

[0023] Determine the target reference monitoring video frame of the target object from the multi-frame monitoring video frames;

[0024] Crop the target reference monitoring video frame based on the area of the target object in the target reference monitoring video frame to obtain the reference monitoring video frame; among them, the reference monitoring video frame contains the area of the target object.

[0025] Among them, the determining whether the target object has a preset behavior by using the sequence of joint point information and the reference monitoring video frame includes:

[0026] Using the joint point information sequence, determine a first confidence level that the target object has the preset behavior;

[0027] Using the reference monitoring video frames, determine a second confidence level that the target object has the preset behavior;

[0028] Perform a weighted sum of the first confidence level and the second confidence level to obtain a comprehensive confidence level that the target object has the preset behavior;

[0029] Using the comprehensive confidence level and a confidence level threshold, determine whether the target object has the preset behavior.

[0030] Another technical solution adopted by this application is to provide a target behavior recognition device, which includes a monitoring module, a joint point module, a reference module, and an identification module; wherein,

[0031] The monitoring module is configured to obtain a tracking video of a target object, and the tracking video includes multiple monitoring video frames containing the target object;

[0032] The joint point module is configured to determine a joint point information sequence of the target object based on the multiple monitoring video frames, and the joint point information sequence includes joint point information corresponding to each monitoring video frame in the multiple monitoring video frames;

[0033] The reference module is configured to determine a reference monitoring video frame of the target object based on the multiple monitoring video frames, and the reference monitoring video frame is obtained based on at least one monitoring video frame in the multiple monitoring video frames;

[0034] The identification module is configured to use the joint point information sequence and the reference monitoring video frame to determine whether the target object has a preset behavior.

[0035] Another technical solution adopted by this application is to provide a target behavior recognition device, which includes a memory and a processor coupled to the memory;

[0036] Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the target behavior recognition method as described above.

[0037] Another technical solution adopted by this application is to provide a computer storage medium, which is used to store program data, and when the program data is executed by a computer, it is used to implement the target behavior recognition method as described above.

[0038] The beneficial effects of the present application are as follows: The target behavior recognition device acquires the tracking video of the target object, and the tracking video includes multiple monitoring video frames containing the target object; based on the multiple monitoring video frames, the joint point information sequence of the target object is determined, and the joint point information sequence includes the joint point information corresponding to each monitoring video frame in the multiple monitoring video frames; and based on the multiple monitoring video frames, the reference monitoring video frame of the target object is determined; by using the joint point information sequence and the reference monitoring video frame, it is determined whether the target object has a preset behavior. The target behavior recognition method of the present application can combine two modalities of information, namely the joint point information sequence and the reference monitoring video frame, to identify the behavior of the target in the monitoring video, improve the recognition accuracy of the target behavior recognition method by using the combined modality information, and speed up the recognition speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a flowchart of an embodiment of the target behavior recognition method provided by the present application;

[0041] Figure 2 It is a flowchart of the automatic control system of the monkey car provided by the present application;

[0042] Figure 3 It is a flowchart of the monkey car passenger detection method based on multi-modal combination provided by the present application;

[0043] Figure 4 It is a schematic structural diagram of an embodiment of the target behavior recognition device provided by the present application;

[0044] Figure 5 It is a schematic structural diagram of another embodiment of the target behavior recognition device provided by the present application

[0045] Figure 6 It is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0047] This application is directed to the detection of people and vehicles in a vehicle control system. By detecting whether there are passengers in the vehicle, it automatically controls the opening and closing of the vehicle. For example, when it detects that there are passengers in the vehicle, it starts the vehicle; when it detects that there are no passengers in the vehicle, it closes the vehicle, which can effectively ensure the normal operation of the vehicle control system, meet the needs of passengers, and also effectively save power resources.

[0048] The target behavior recognition method provided by this application can be applied to the detection of people and vehicles in multiple fields, such as roller coasters in amusement parks, monkey cars in mines, and shuttle buses in scenic spots. In the embodiments of this application, taking the detection of monkey cars in mines as a specific implementation method, there are no restrictions on other application fields.

[0049] The main problem solved by this application is to achieve the detection of whether there are people on the monkey car at the camera device end. For the behavior of whether a person is sitting on the monkey car, this application combines two modalities of information, namely joint point sequence information and image information, and uses the methods of multi-modal information fusion and multi-camera alarm information fusion to accurately control the operation and stop of the monkey car.

[0050] For details, please refer to Figure 1 and Figure 2 , Figure 1 which are the schematic flowcharts of an embodiment of the target behavior recognition method provided by this application, Figure 2 and

[0051] As Figure 1 shown, the target behavior recognition method of the embodiments of this application can specifically include the following steps:

[0052] Step S11: Obtain the tracking video of the target object, where the tracking video includes multiple monitored video frames containing the target object.

[0053] In the embodiments of this application, as Figure 2 shown, multiple camera devices are set in the mine tunnel in the underground mine. The target behavior recognition device uses one of the multiple camera devices, that is, the first camera, to collect monitored video data of the area in the mine tunnel, and then processes the monitored video data through the PLC (programmable logic controller) of the monkey car control system to obtain the tracking video of the target object, that is, the monkey car riding video.

[0054] It should be noted that in the underground mine scenario, the target object can be the staff in the mine tunnel. The target behavior recognition method of the embodiments of this application can also be applied to other application scenarios, and specific application scenarios and their target objects are not listed one by one here.

[0055] Further, the target behavior recognition device extracts multiple monitored video frames containing the target object and the temporal relationship of the multiple monitored video frames from the tracking video. It should be noted that the multiple monitored video frames can be multiple consecutive monitored video frames in terms of time sequence, or multiple monitored video frames respectively extracted at different temporal positions according to the time sequence of the tracking video to form the multiple monitored video frames.

[0056] Please continue to refer to Figure 3 , Figure 3 which is a schematic flowchart of the monkey cart passenger detection method based on multimodal combination provided by this application. As Figure 3 shown, the target behavior recognition device inputs the tracking video, that is, the multiple monitored video frames, into the human body recognition and tracking module for target tracking in the tracking video. Here, the target referred to in the embodiments of this application is generally a pedestrian or a passenger.

[0057] The human body recognition and tracking module generates a target detection box on each monitored video frame by performing target detection on each monitored video frame. Furthermore, the same identifier is set for the target detection boxes of the same target object in the multiple monitored video frames to distinguish the multiple targets within the multiple monitored video frames.

[0058] Specifically, the human body recognition unit in the human body recognition and tracking module identifies the human body from each monitored video frame, and in combination with the tracking unit, saves each target person in the tracking video according to different ID (Identity document) information to generate image modality information, and sends it to the joint point extraction module to be converted into joint point modality information. Then, the image modality information is sent to the picture-based behavior recognition module to judge the human behavior, and the joint point modality information is sent to the joint point sequence behavior recognition module to judge the human behavior. Finally, the result analysis module combines the output results of the two modalities to generate and output the final result.

[0059] Among them, the human body recognition and tracking module includes a human body recognition unit and a tracking unit. The human body recognition unit is used to detect the human body box, that is, the target detection box. The improved YOLO V3 network is used in the human body recognition unit. The improved YOLO V3 network means that its original backbone network is replaced with a residual network, and the others remain unchanged. In the human body recognition unit, the improved YOLO V3 network uses various human targets as the training set to train the network. When in use, the improved YOLO V3 network identifies the position of the human body box through logistic regression, and expands the human body box output by the network by a factor of two centered on itself to generate the final target detection box.

[0060] The tracking unit uses the Camshift algorithm (continuously adaptive MeanShift algorithm) to track the recognized human target, so that the detected human target uses the same ID in multiple frames of surveillance video frames. Among them, the Camshift algorithm can track the moving target object in multiple frames of surveillance video frames. Based on the color information of the moving object in multiple frames of surveillance video frames as features, perform Mean-Shift operations on each frame of the input multiple frames of surveillance video frames respectively, and use the target center and search window size (kernel function bandwidth) of the previous frame of surveillance video frame as the initial values of the center and search window size of the Mean-Shift algorithm for the next frame of surveillance video frame. Iterating like this can achieve tracking of the position of the target in multiple frames of surveillance video frames.

[0061] It should be noted that the specific method used by the human body recognition and tracking module is not limited in the embodiments of the present application. That is, in other embodiments, other object detection networks and / or other object tracking algorithms can also be used.

[0062] It should be noted that Figure 3 The human body recognition and tracking module, joint point extraction module, behavior recognition module based on joint point sequence, behavior recognition module based on pictures, and result analysis module shown can all be regarded as component modules of the target behavior recognition device, or as component modules in the behavior recognition system carried on the target behavior recognition device.

[0063] Step S12: Based on multiple frames of surveillance video frames, determine the joint point information sequence of the target object. The joint point information sequence includes the joint point information corresponding to each surveillance video frame in multiple frames of surveillance video frames.

[0064] In the embodiments of the present application, as Figure 3 shown, the target behavior recognition device respectively inputs the target tracking information obtained by the human body recognition and tracking module, such as the target detection box, into the joint point extraction module and the behavior recognition module based on pictures.

[0065] The joint point extraction module selects the corresponding target according to the identifier set in step S11, extracts the target detection box of the target in multiple frames of surveillance video frames, and thus extracts the joint point information of the target object by using multiple frames of surveillance video frames. Among them, the joint point information at least includes the joint point position coordinates of the target object, etc.

[0066] Specifically, the joint point extraction module can first extract the joint point features of the target object in each frame of the surveillance video frame, and then utilize the temporal relationship of multiple frames of the surveillance video frame to sequentially correlate and fuse the joint point features of multiple frames of the surveillance video frame, thereby obtaining the joint point information of the target in multiple consecutive frames of the surveillance video frame. Among them, the joint point information also includes the spatio-temporal joint point features of the target, that is, the spatial features of the joint points of the target object within the image and the temporal features within the image sequence can be reflected.

[0067] The joint point sequence information refers to the joint point sequence information of each person, including the spatial information and temporal information of the joint points. The spatial information of the joint points includes the two-dimensional coordinate positions of 13 joint points of the human body, and the temporal information of the joint points includes the temporal information of 30 consecutive frames. The spatio-temporal graph convolution recognition unit uses the spatio-temporal graph convolution model to recognize the joint point sequence. The spatio-temporal graph convolution model includes temporal convolution and graph convolution. Temporal convolution is used to learn the temporal information between joint point sequences, and graph convolution is used to learn the spatial information between joint points.

[0068] Furthermore, the joint point extraction module can convert the human body information within the target detection box in the surveillance video frame into the joint point modality. In this application, the High-Resolution Network (HRNet) is used. HRNet forms more stages by gradually increasing the high-resolution to low-resolution subnets, and connects the multi-resolution subnets in parallel. It performs multi-scale repeated fusion by repeatedly exchanging information on the parallel multi-resolution subnets, and finally estimates the human body joint points through the network output.

[0069] It should be noted that for the specific method used by the joint point extraction module, there is no limitation in the embodiments of this application. That is, in other embodiments, other feature extraction networks can also be used to extract relevant joint point information.

[0070] Step S13: Based on multiple frames of the surveillance video frame, determine the reference surveillance video frame of the target object. The reference surveillance video frame is obtained based on at least one frame of the surveillance video frame among multiple frames of the surveillance video frame.

[0071] In the embodiments of this application, since the behavior recognition network in the behavior recognition module based on pictures can output the behavior recognition result of the target object for only a single picture, the behavior recognition module based on pictures can also select at least one frame of the surveillance video frame from multiple frames of the surveillance video frame as the reference surveillance video frame and input it into the behavior recognition module based on pictures for the behavior recognition of the target object. Among them, the reference surveillance video frame can be a randomly selected frame from multiple frames of the surveillance video frame, or a frame with the best image quality, such as the highest resolution, or a frame at a fixed position in multiple frames of the surveillance video frame, such as the middle frame, etc. Only inputting a single frame image into the behavior recognition module based on pictures can effectively reduce the time consumption of network behavior recognition and improve the behavior recognition efficiency.

[0072] Further, the target behavior recognition device can determine at least one target reference surveillance video frame of the target object from multiple frames of surveillance video frames in the above manner. Then, based on the region of the target object in the target reference surveillance video frame, the target reference surveillance video frame is cropped to obtain a reference surveillance video frame. The reference surveillance video frame at least includes the region of the target object. Through this cropping method, the amount of information input to the picture behavior recognition module can be effectively reduced, and the speed of picture behavior recognition can be increased.

[0073] Step S14: Using the joint point information sequence and the reference surveillance video frame, determine whether the target object has a preset behavior.

[0074] In the embodiment of the present application, on the one hand, the target behavior recognition device uses the joint point information sequence to determine the sequence action information of the target object, and on the other hand, uses the reference surveillance video frame to determine the behavior reference information of the target object.

[0075] Using the joint point information sequence to determine the sequence action information of the target object is specifically as follows:

[0076] Specifically, in the embodiment of the present application, the spatio-temporal graph data of the target object is generated based on the joint point information of the target object extracted by the joint point extraction module by the joint point sequence behavior recognition module, that is, the time feature and the space feature of the joints of the target are fused in the form of graph data. Then, the joint point sequence behavior recognition module processes the spatio-temporal graph data of the target using a graph convolutional neural network, and based on the spatio-temporal graph data of the target object, outputs the predicted sequence action and confidence of the target object in multiple frames of surveillance video frames, that is, generates the final sequence action information.

[0077] It should be noted that the sequence action information of the target object may include multiple preset sequence actions and the confidence corresponding to each preset sequence action; it may also only include the preset sequence action with the highest confidence, which is not limited here.

[0078] For example, determining the preset sequence action that matches the joint point information sequence among multiple preset sequence actions as the sequence action information of the target object can be regarded as one or more preset sequence actions with a confidence higher than the preset confidence threshold being the matching preset sequence actions.

[0079] It should be noted that in the application scenario of an underground mine, the multiple frames of surveillance video frames in the above steps are collected for the moving tools in the underground mine, and the transportation tools include but are not limited to tools with steel cables as tracks, such as man-cages. Based on this, the multiple frames of surveillance video frames can also specifically be collected for the area where the target object enters or leaves the transportation tool, such as the area where the staff or passengers enter or leave the man-cage as the surveillance area.

[0080] Specifically, the joint point sequence-based behavior recognition module mainly uses the spatiotemporal graph convolution recognition unit and joint point sequence information to determine whether a person is sitting on the monkey car or other behaviors. The spatiotemporal graph convolution recognition unit is composed of a spatiotemporal graph convolution network and can recognize any customized sequence actions.

[0081] The custom sequence actions of this application include at least: a person sitting on a monkey car and a person not sitting on a monkey car. Further, the custom sequence actions may also include: a person sitting on a monkey car (the monkey car is running), a person sitting on a monkey car (the monkey car is not running), a person walking in a mine tunnel, a person running in a mine tunnel, and a person standing still in a mine tunnel. By collecting custom data, the spatiotemporal graph convolutional network is trained. When the indicators meet the requirements, the model is deployed on the camera in the mine tunnel. Among them, the operation formula of the spatiotemporal graph convolutional network is as follows:

[0082]

[0083] in, is the “graph convolution kernel”, X is the joint point sequence, X′ is the updated joint point sequence, k is the “number of graph convolution kernels”, n represents BatchSize (the number of samples selected for one training), c represents the features of the joint points (i, y), t represents the length of the joint point sequence (i.e. the number of images), and v and w represent the spatial features of the joint points.

[0084] Finally, based on the joint point sequence behavior recognition module, the updated target sequence action features are compared with the action features of multiple preset sequence actions, and the preset sequence action with the highest feature similarity is determined as the sequence action information of the target object.

[0085] Using the reference surveillance video frame, the behavior reference information of the target object is determined as follows:

[0086] In the embodiment of the present application, the picture-based behavior recognition module is a classification model designed with Xception as the backbone network, which can directly identify the action category based on the color image of the target person. Among them, the training process of the action classification model refers to first collecting the behavior pictures that need to be identified as the training set. The training set of the present application can only contain two types of data: riding a monkey car and not riding a monkey car. Then send it to the Xception network for training. If the Xception network reaches the predetermined goal on the test set, the training stops and the network parameters are saved; in the recognition process, the color image of the target person that needs to be identified refers to the intermediate frame image corresponding to the joint point sequence, that is, the picture-based behavior recognition module judges the target person's action in the video once every 30 frames; the action classification result recognized by the final model, that is, the behavior reference information is directly sent to the result analysis module.

[0087] It should be noted that in other embodiments, other behavior recognition networks can also be used to recognize the target behavior in the image, which will not be enumerated one by one here.

[0088] Furthermore, the result analysis module can perform a weighted sum of the outputs of the two modal information of the tracking video, that is, the sequence action information and the behavior reference information. The weighted sum formula for the output of the two modal information by a single camera is as follows:

[0089] R = αR ST-GCN + βR pic

[0090] where R is the final output result, α and β are both weight coefficients, and α + β = 1; R ST-GCN is the result based on the joint point sequence behavior recognition and module, and R pic is the recognition result based on the picture behavior.

[0091] It should be noted that the result based on the joint point sequence behavior recognition and module is the confidence level of determining that the target object has a preset behavior using the sequence action information; the recognition result based on the picture behavior is the confidence level of determining that the target object has a preset behavior using the reference monitoring video frame.

[0092] By performing a weighted sum of the two modal information output by a single camera, the comprehensive behavior recognition result of the single camera for the target can be obtained. It should be noted that the preset behavior used for the weighted sum should be the same sequence action, and the result obtained by the weighted sum is the final confidence level of the sequence action. The result analysis module uses the sequence action with the highest final confidence level as the sequence action finally recognized and output by the target in the monitoring video, and then determines whether the target object has a preset behavior. For example, the target behavior recognition device can obtain the final confidence level of the preset behavior, and then compare the final confidence level with the confidence threshold. If the confidence level of the preset behavior is higher than or equal to the confidence threshold, it is determined that the target object has the preset behavior; if the confidence level of the preset behavior is lower than the confidence threshold, it is determined that the target object does not have the preset behavior.

[0093] Further, as shown in step 14, since the joint point sequence behavior recognition module can output more custom sequence actions by detecting joint points than the image-based behavior recognition module, the custom sequence actions output by the joint point sequence behavior recognition module can be used to supplement and assist the sequence actions of the final recognition output. For example, the action of the sequence finally recognized by the result analysis module is that a person is sitting on a man-riding vehicle. The actions of a person sitting on a man-riding vehicle (the man-riding vehicle is running) and a person sitting on a man-riding vehicle (the man-riding vehicle is not running) output by the joint point sequence behavior recognition module can be used to supplement and describe the specific sequence actions of a person sitting on a man-riding vehicle. For example, when a person is sitting on a man-riding vehicle and the man-riding vehicle is running, the current state of the man-riding vehicle running system can be maintained; when a person is sitting on a man-riding vehicle and the man-riding vehicle is not running, the man-riding vehicle running system needs to be started.

[0094] Finally, the target behavior recognition device can perform corresponding preset operations according to the final sequence actions. For example, when the finally output sequence action is that a person is sitting on a man-riding vehicle, the man-riding vehicle running system is run; when the finally output sequence action is that a person is not sitting on a man-riding vehicle, the man-riding vehicle is turned off.

[0095] Further, the result analysis module can also perform weighted summation on the analysis of multiple camera results. The analysis of multiple camera results means that if all cameras do not output a signal of riding on a man-riding vehicle, the man-riding vehicle running system is turned off; if any one or more cameras detect that someone is riding on a man-riding vehicle, the man-riding vehicle running system is run.

[0096] In the embodiment of the present application, the target behavior recognition device acquires a tracking video of a target object, and the tracking video includes multiple monitoring video frames containing the target object; based on the multiple monitoring video frames, a joint point information sequence of the target object is determined, and the joint point information sequence includes the joint point information corresponding to each monitoring video frame in the multiple monitoring video frames; and based on the multiple monitoring video frames, a reference monitoring video frame of the target object is determined; using the joint point information sequence and the reference monitoring video frame, it is determined whether the target object has a preset behavior. The target behavior recognition method of the present application can combine two modalities of information, namely the joint point information sequence and the reference monitoring video frame, to recognize the behavior of the target in the monitoring video, and improve the recognition accuracy of the target behavior recognition method and speed up the recognition speed by using the combined modality information.

[0097] The above embodiments are only one common case of the present application, and do not impose any limitation on the technical scope of the present application. Therefore, any minor modifications, equivalent changes or decorations made to the above content based on the essence of the solution of the present application still fall within the scope of the technical solution of the present application.

[0098] Please continue to refer to Figure 4 , Figure 4It is a schematic structural diagram of an embodiment of the target behavior recognition device provided by this application. The target behavior recognition device 400 in the embodiment of this application includes a monitoring module 41, a joint point module 42, a reference module 43, and an identification module 44.

[0099] Among them, the monitoring module 41 is used to obtain the tracking video of the target object, and the tracking video includes multiple monitoring video frames containing the target object.

[0100] The joint point module 42 is used to determine the joint point information sequence of the target object based on the multiple monitoring video frames, and the joint point information sequence includes the joint point information corresponding to each monitoring video frame in the multiple monitoring video frames.

[0101] The reference module 43 is used to determine the reference monitoring video frame of the target object based on the multiple monitoring video frames, and the reference monitoring video frame is obtained based on at least one monitoring video frame in the multiple monitoring video frames.

[0102] The identification module 44 is used to determine whether the target object has a preset behavior by using the joint point information sequence and the reference monitoring video frame.

[0103] Please continue to refer to Figure 5 , Figure 5 It is a schematic structural diagram of another embodiment of the target behavior recognition device provided by this application. The target behavior recognition device 500 in the embodiment of this application includes a processor 51, a memory 52, an input / output device 53, and a bus 54.

[0104] The processor 51, the memory 52, and the input / output device 53 are respectively connected to the bus 54. Program data is stored in the memory 52, and the processor 51 is used to execute the program data to implement the target behavior recognition method described in the above embodiment.

[0105] In the embodiment of this application, the processor 51 can also be called a CPU (Central Processing Unit, central processing unit). The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, FieldProgrammable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 51 can also be any conventional processor, etc.

[0106] The present application also provides a computer storage medium. Please continue to refer to Figure 6 , Figure 6 FIG. 6 is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. Program data 61 is stored in the computer storage medium 600. When the program data 61 is executed by a processor, it is used to implement the target behavior recognition method of the above embodiment.

[0107] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0108] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. A method for identifying a target behavior, characterized in that, The described target behavior recognition method includes: Obtaining a tracking video of a target object, where the tracking video includes multiple monitored video frames containing the target object; Based on the multiple monitored video frames, determining a sequence of joint point information of the target object, where the sequence of joint point information includes the joint point information corresponding to each monitored video frame in the multiple monitored video frames; and Based on the multiple monitored video frames, determining a reference monitored video frame of the target object, where the reference monitored video frame is obtained based on at least one of the multiple monitored video frames; Using the sequence of joint point information and the reference monitored video frame to determine whether the target object has a preset behavior; Among them, using the sequence of joint point information and the reference monitored video frame to determine whether the target object has a preset behavior includes: Based on the sequence of joint point information, obtaining the target sequence action features of the target object; Using a spatio-temporal graph convolutional network to update the target sequence action features; Comparing the updated target sequence action features with the action features of multiple preset sequence actions respectively, and determining the preset sequence action with the highest feature similarity as the sequence action information of the target object; Using the reference monitored video frame to determine the behavior reference information of the target object; Based on the sequence action information and the behavior reference information, determining whether the target object has the preset behavior.

2. The target behavior recognition method according to claim 1, wherein The multiple monitored video frames are collected for a transportation tool; the sequence action information is determined based on at least one of the position distribution information of the target object and the transportation tool, and the posture of the target object.

3. The target behavior recognition method according to claim 2, wherein The transportation tool includes a tool with a steel cable as a track, and the multiple monitored video frames are collected for the area where the target object enters or leaves the transportation tool.

4. The target behavior recognition method according to claim 1, wherein The determining the reference monitored video frame of the target object based on the multiple monitored video frames includes: Determining a target reference monitored video frame of the target object from the multiple monitored video frames; Based on the area of the target object in the target reference monitored video frame, cropping the target reference monitored video frame to obtain the reference monitored video frame; where the reference monitored video frame contains the area of the target object.

5. The target behavior recognition method according to claim 1, wherein Using the sequence of joint point information and the reference monitored video frame to determine whether the target object has a preset behavior includes: Using the sequence of joint point information to determine a first confidence level that the target object has the preset behavior; Using the reference monitored video frame to determine a second confidence level that the target object has the preset behavior; Weighted summing the first confidence level and the second confidence level to obtain a comprehensive confidence level that the target object has the preset behavior; Using the comprehensive confidence level and a confidence level threshold to determine whether the target object has the preset behavior.

6. An apparatus for recognizing a target behavior, characterized in that, The target behavior recognition device includes a monitoring module, a joint point module, a reference module, and an identification module; where The monitoring module is used to obtain the tracking video of the target object, and the tracking video includes multiple monitoring video frames containing the target object; The joint point module is used to determine the joint point information sequence of the target object based on the multiple monitoring video frames, and the joint point information sequence includes the joint point information corresponding to each monitoring video frame in the multiple monitoring video frames; The reference module is used to determine the reference monitoring video frame of the target object based on the multiple monitoring video frames, and the reference monitoring video frame is obtained based on at least one monitoring video frame in the multiple monitoring video frames; The recognition module is used to determine whether the target object has a preset behavior by using the joint point information sequence and the reference monitoring video frame; Among them, the recognition module is specifically used to obtain the target sequence action feature of the target object based on the joint point information sequence; update the target sequence action feature by using the spatio-temporal graph convolutional network; compare the updated target sequence action feature with the action features of multiple preset sequence actions respectively, and determine the preset sequence action with the highest feature similarity as the sequence action information of the target object; determine the behavior reference information of the target object by using the reference monitoring video frame; and determine whether the target object has the preset behavior based on the sequence action information and the behavior reference information.

7. An apparatus for recognizing a target behavior, characterized in that, The target behavior recognition device includes a memory and a processor coupled to the memory; Among them, the memory is used to store program data, and the processor is used to execute the program data to implement the target behavior recognition method according to any one of claims 1 to 5.

8. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the target behavior recognition method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Vehicle control method, vehicle control device and computer readable storage medium

    CN110239529A

  • Behavior detection method and device, computer equipment and storage medium

    CN113657155A