Abnormal behavior recognition method, medium, product, electronic equipment and vehicle
By fusing the dynamic and static characteristics of multi-frame images, the problem of insensitive response of DMS systems when key target occlusion or angle changes is solved, and the accuracy of abnormal behavior recognition and system real-time performance is improved.
Patent Information
- Application Number
- CN202510404381.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-12
AI Technical Summary
When existing DMS systems face key targets being occluded or angular changes, their responses are not sensitive enough, resulting in a decrease in detection accuracy and reliability. Multi-frame image feature extraction increases the calculation amount, affecting the real-time performance of the system.
By fusing the dynamic and static features of the object detection object in the multi-frame image without increasing the calculation amount, including detection frames and key point information, abnormal behavior is identified.
It improves the accuracy and reliability of abnormal behavior recognition, reduces the misclassification of dangerous behaviors, and ensures the real-time and computing efficiency of the system.
Smart Images

Figure CN120472419A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle technology, and in particular to an abnormal behavior identification method, medium, product, electronic device and vehicle. Background Art
[0002] With the rapid development of intelligent driving technology, driver monitoring systems (DMS) have become a key technology in vehicle safety. Related technologies typically analyze image features in a single frame, resulting in insensitive responses to occlusion or changes in the angle of key targets, reducing the accuracy and reliability of object monitoring. Summary of the Invention
[0003] The embodiments of the present application provide an abnormal behavior identification method, medium, product, electronic device and vehicle to solve the above-mentioned problems.
[0004] To achieve the above objectives, according to a first aspect of the present application, a method for identifying abnormal behavior is provided, the method comprising:
[0005] Based on the collected images of the detection area, the dynamic characteristics of the target detection object are obtained; the detection area is the area where abnormal behavior is observed;
[0006] Abnormal behavior is identified based on the dynamic characteristics of the target detection object.
[0007] Optionally, obtaining dynamic features of the target detection object based on the captured image of the detection area includes:
[0008] Based on the collected image of the detection area, obtaining target information in the collected image;
[0009] Based on the target information, dynamic features of the target detection object are obtained.
[0010] Optionally, the target information includes a detection box and / or key points.
[0011] Optionally, obtaining dynamic features of the target detection object based on the target information includes:
[0012] Based on the target information, determining a target key point in the acquired image corresponding to a target part of the target detection object;
[0013] Based on the target key point information in each frame of the collected image, the motion information of the target part is obtained to determine the dynamic feature.
[0014] Optionally, the target key point information includes position information of the target key point.
[0015] Optionally, obtaining the motion information of the target part based on the target key point information in each frame of the acquired image includes:
[0016] Based on the position change information of the target key point in the multiple frames of the acquired image, the motion information of the target part is obtained.
[0017] Optionally, obtaining the motion information of the target part based on the position change information of the target key point in the multiple frames of the acquired image includes:
[0018] First motion information is obtained based on position change information of the target key point in a group of key frame images within a first time interval.
[0019] Optionally, obtaining the motion information of the target part based on the position change information of the target key point in the multiple frames of the acquired image further includes:
[0020] Second motion information is obtained based on position change information of the target key point in a group of key frame images at a second time interval, where the first time interval is smaller than the second time interval.
[0021] Optionally, the method further includes:
[0022] The dynamic feature is determined based on the motion information of the target part and the key point distance information of the target part.
[0023] Optionally, the key point distance information is obtained by:
[0024] Based on the distance between every two target key points of the target part, key point distance information of the target part is determined.
[0025] Optionally, the method further includes:
[0026] Based on the collected images of the detection area, the static features of the target detection object are obtained;
[0027] The abnormal behavior is identified based on the dynamic features and the static features of the target detection object.
[0028] Optionally, the static features include image features of a target area of the target detection object in the captured image.
[0029] The method further comprises:
[0030] Determine, based on multiple detection frames in the captured image, a detection frame closest to the target detection frame;
[0031] Based on the detection frame closest to the target detection frame, image features of the target area are determined, wherein the target area includes at least one target part of the target detection object.
[0032] Optionally, determining a detection frame closest to a target detection frame based on multiple detection frames in the captured image includes:
[0033] Based on the distance between the target detection frame and other detection frames and a preset distance threshold, the detection frame closest to the target detection frame is determined.
[0034] Optionally, determining the image features of the target area based on the detection frame closest to the target detection frame includes:
[0035] Merge the detection frame closest to the target detection frame with the target detection frame to obtain a circumscribed rectangular frame;
[0036] Feature extraction is performed on the image area corresponding to the circumscribed rectangular frame to obtain image features of the target area of the target detection object.
[0037] Optionally, the target area is one of a face area, a hand area, and a whole body area of the target detection object.
[0038] Optionally, the target area is a facial area of the target detection object, and the target detection frame is a facial detection frame.
[0039] Optionally, the static feature includes a relative position feature,
[0040] The method further comprises:
[0041] Based on the plurality of detection frames in the captured image, calculating the distance between each two detection frames to determine a detection frame distance matrix;
[0042] Feature extraction is performed on the detection frame distance matrix to obtain the relative position feature.
[0043] Optionally, identifying abnormal behavior includes:
[0044] Feature fusion is performed based on static features and dynamic features to determine the behavior recognition result of the target detection object.
[0045] Optionally, the performing feature fusion based on static features and dynamic features to determine the behavior recognition result of the target detection object includes:
[0046] Based on the first weight and the second weight, the dynamic feature and the static feature are weightedly fused to obtain a fused feature;
[0047] Feature extraction is performed on the fused features to determine a behavior recognition result of the target detection object.
[0048] Optionally, the detection frame includes a detection frame of at least one of a face, a hand, and an abnormal object related to abnormal behavior of the target detection object.
[0049] Optionally, the detection area is the driving area of the vehicle, and the target detection object is the driver.
[0050] According to the second aspect of the present application, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to implement any one of the abnormal behavior identification methods provided in the embodiments of the present application.
[0051] According to the third aspect of the present application, an embodiment of the present application further provides a computer program product, which stores instructions, and when the instructions are executed by a computer, enables the computer to implement any one of the abnormal behavior identification methods provided in the embodiments of the present application.
[0052] According to a fourth aspect of the present application, an embodiment of the present application further provides an electronic device, including:
[0053] a memory having a computer program stored thereon;
[0054] A processor is used to execute the computer program in the memory to implement any of the abnormal behavior identification methods provided in the embodiments of the present application.
[0055] According to the fifth aspect of the present application, an embodiment of the present application also provides a vehicle, including the electronic device described above, or the abnormal behavior recognition device described above.
[0056] Some embodiments of the present specification include at least the following beneficial effects: by determining dynamic features, the extraction of multiple features of the target detection object is achieved without consuming a large amount of computation, which makes up for the problem that the related technology only extracts the image features of the current frame, resulting in the occlusion of the key targets of the target detection object and the insensitivity to the angle changes of the key targets, as well as the problem of excessive computation caused by extracting the image features of multiple frames of images. Without increasing the computational burden, the dynamic features of the target detection object are introduced into the algorithm to better classify behaviors, reduce more dangerous behaviors of the target detection object, and ensure safety as much as possible.
[0057] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0059] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.
[0060] Figure 1 is an application scenario diagram of the abnormal behavior identification method according to some embodiments of this specification;
[0061] Figure 2 is an exemplary schematic diagram of an abnormal behavior identification system according to some embodiments of this specification;
[0062] Figure 3 is an exemplary flow chart of a method for identifying abnormal behavior according to some embodiments of this specification;
[0063] Figure 4 is an exemplary schematic diagram of determining motion information according to some embodiments of this specification;
[0064] Figure 5 is an exemplary schematic diagram of image features of a target area of a target detection object according to some embodiments of this specification;
[0065] Figure 6 is an exemplary flow chart of relative position features according to some embodiments of this specification;
[0066] Figure 7 is an exemplary schematic diagram of weighted fusion according to some embodiments of this specification;
[0067] Figure 8 is a schematic structural diagram of an electronic device according to some embodiments of this specification;
[0068] Figure 9 is an exemplary schematic diagram of a vehicle according to some embodiments of the present specification. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the technical target detection object in this field without paying creative labor are within the scope of protection of this application.
[0070] In order to facilitate understanding of the implementation scheme provided in the embodiment of the present application, the relevant application background of the abnormal behavior identification method provided in the embodiment of the present application is first explained.
[0071] Currently, most existing DMS systems only extract image features from the current frame. This single-frame feature extraction approach is susceptible to insufficient sensitivity when key targets are obscured or their angles change. For example, if a driver's face or hands are obscured, the DMS system may be unable to accurately identify dangerous movements, affecting detection accuracy and reliability.
[0072] To compensate for the shortcomings of single-frame feature extraction, some algorithms attempt to extract features from multiple frames to obtain more comprehensive motion information. However, this approach increases computational complexity, placing an excessive burden on the vehicle's computer system. For example, frequently performing complex feature extraction and analysis on multiple frames consumes significant computing resources, impacting the system's real-time performance and responsiveness.
[0073] In view of this, some embodiments of the present specification provide a method for identifying abnormal behavior, which extracts the position and size information of the detection frame of the target detection object and the position and size information of the detection frame of the abnormal object, and fuses and classifies the information with the key point information in multiple frames of captured images (e.g., the previous frame and the subsequent frame of the current frame), and outputs a final target detection object behavior recognition result. The DMS system includes target information for identifying the gestures and head posture of the target detection object, so that the acquisition of the key points of the hand and face of the target detection object does not consume additional computational resources; moreover, the algorithm used for feature extraction of the key points of the hand and face in the multiple frames of captured images and the algorithm used for feature extraction of the relative position of the detection frame related to the abnormal behavior both consume a small amount of computational resources, while enhancing the algorithm's feature attention ability, avoiding increasing the algorithm's computational resources and power consumption. In the case where the object related to the dangerous action of the target detection object is occluded or the detection fails, the correct target detection object behavior recognition result can still be output in combination with the information of the multiple frames of captured images, which helps to refine the behavior classification of the abnormal behavior of the target detection object, thereby realizing different prompting methods according to different abnormal behaviors, and further improving the reliability and accuracy of monitoring the target detection object.
[0074] Figure 1 This is an application scenario diagram of the abnormal behavior identification method shown in some embodiments of this specification.
[0075] like Figure 1 As shown, the application scenario 100 of the abnormal behavior recognition method may include a processing device 110 , a network 120 , a vehicle 130 , and an image acquisition device 140 .
[0076] The processing device 110 can be a local physical server, a cloud server (such as Figure 1 as shown), a virtual server, a distributed server, or any other suitable computing device.
[0077] In some embodiments, processing device 110 may be integrated into vehicle 130 .
[0078] In some embodiments, processing device 110 may acquire image data of a target detection object within the vehicle. Processing device 110 may use the acquired data to determine abnormal behavior of the target detection object. For example, processing device 110 may obtain data from image acquisition device 140. Processing device 110 may communicate with image acquisition device 140 and / or other components of vehicle 130 via network 120.
[0079] The network 120 may include but is not limited to any one or combination of wireless local area network (WLAN), wide area network (WAN), wireless network of radio wave / cellular network / satellite communication network and / or local wireless network or short-range wireless network (Bluetooth™).
[0080] Vehicle 130 may be a transportation service vehicle for application scenario 100. In some embodiments, vehicle 130 may include a taxi, a private car, a ride-sharing vehicle, a shared vehicle, a bus, a bicycle, a tricycle, a motorcycle, a robot, an AGV, or any combination thereof. In some embodiments, vehicle 130 may include a private self-driving car, a shared self-driving car, an unmanned automatic freight car, or the like. In some embodiments, vehicle 130 may receive instructions issued by processing device 110 and perform corresponding tasks in accordance with the instructions. For example, vehicle 130 may receive instructions such as executing an autonomous driving operation and automatically control vehicle 130 to complete the corresponding operation.
[0081] Image acquisition device 140 is used to capture image data of the target detection object. In some embodiments, image acquisition device 140 includes, but is not limited to, a camera, a webcam, a camcorder, or other devices with a camera function (such as a mobile phone or tablet). Image acquisition device 140 can record a video consisting of multiple image frames captured at multiple time points. In some embodiments, image acquisition device 140 can capture a series of image frames within the vehicle. The image frames can be sent to processing device 110 in real time (e.g., via streaming) or based on a preset period.
[0082] In some embodiments, the image acquisition device 140 can be mounted on a vehicle's steering wheel, dashboard, or A-pillar. In some embodiments, the mounting structure can use screws, adhesives, or other mounting structures. In some embodiments, any suitable mounting structure can also be used.
[0083] In some embodiments, the image acquisition device 140 can be deployed at least one location within the vehicle. In some embodiments, the number and locations of the image acquisition devices 140 can be predetermined. For example, the vehicle interior environment can be divided into multiple areas, each of which can be equipped with an image acquisition device 140. In some embodiments, the image acquisition device 140 can communicate with one or more components in the application scenario (e.g., the vehicle 130 or the processing device 110) via the network 120.
[0084] It should be noted that application scenario 100 is provided for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can make various modifications or variations based on the description of this specification. For example, application scenario 100 can also include a database, an information source, etc. For another example, application scenario 100 can be implemented on other devices to achieve similar or different functions. However, changes and modifications will not deviate from the scope of this specification.
[0085] Figure 2 is an exemplary schematic diagram of an abnormal behavior identification system according to some embodiments of this specification.
[0086] In some embodiments, the abnormal behavior identification system communicates with a hardware abstraction layer, an application, and the like.
[0087] Among them, the hardware abstraction layer is used to encapsulate the hardware details of image acquisition devices from different manufacturers and provide a unified API interface to upper-level applications.
[0088] In some embodiments, the application interacts with the hardware abstraction layer through the image acquisition device interface in the abnormal behavior recognition system, sends a request (such as taking a photo, recording a video), and the hardware abstraction layer accepts the request and calls the underlying driver to operate the hardware, thereby realizing communication with the image acquisition device and other hardware, and completing the acquisition and output of image data (such as the captured image of the target detection object).
[0089] In some embodiments, the abnormal behavior recognition system can obtain target information of the target detection object based on the collected image, such as the detection frame and key points of the target object of the target detection object through key point detection algorithm, target detection algorithm, etc.
[0090] In some embodiments, the abnormal behavior recognition system can extract and fuse features based on the target information of the target detection object to obtain an output signal, such as the behavior recognition result of the target detection object. For more details about this embodiment, please refer to the relevant description below.
[0091] It should be noted that multiple detection frames and / or key point information associated with abnormal movements of the target detection object can be obtained through the DMS (Driver Monitoring System) without increasing the burden of calculation. The above key point information and detection frames can also be reused by other modules (such as other modules in the DMS).
[0092] Among them, the DMS (Driver Monitoring System) application business module can issue an early warning or take corresponding intervention measures based on the output signal. For example, when abnormal behavior of the target detection object is detected, a sound or vibration reminder is issued to prompt the target detection object. For another example, when the target detection object is detected to be distracted, a warning prompt is issued to guide it to concentrate its attention. For another example, in extreme cases (such as when the target detection object loses consciousness), automatic braking or autonomous driving takeover function is triggered.
[0093] Figure 3 is an exemplary flow chart of the abnormal behavior identification method according to some embodiments of this specification. In some embodiments, process 300 can be executed based on an electronic device. Figure 3 As shown, the process 300 includes the following steps.
[0094] Step 310: obtaining dynamic features of the target detection object based on the captured image of the detection area; the detection area is the area where abnormal behavior is observed.
[0095] In some embodiments, the detection area refers to the area including the target detection object. For example, the detection area can be a specific environmental area, whose definition and scope depend on the application scenario and task requirements, such as an entrance, exit, corridor, parking lot, etc.
[0096] In some embodiments, the detection area is the driving area of the vehicle, and the target detection object is the driver.
[0097] In some embodiments, the processor can be in communication with the image acquisition device to dynamically acquire captured images. For example, the target detection object can be periodically captured while the vehicle is traveling to acquire captured images.
[0098] The captured image may be an image directly captured by the image capture device, or may be an image captured by a method such as frame extraction from an image video captured by the image capture device.
[0099] In some embodiments, each frame of captured image may be a real-time image. For example, the processor may acquire an image captured by the image capture device in real time (e.g., every 0.02 seconds, 0.01 seconds, etc.). In some embodiments, each frame of captured image may also be an image stored in a storage device.
[0100] In some embodiments, the processor may periodically acquire an image video captured by the image capture device, wherein the image video is a sequence of multiple frames of captured images in a time sequence of scanning. In some embodiments, each frame of the captured image in the image video may be temporally continuous.
[0101] In some embodiments, the processor may capture images of each frame in the image video in chronological order and execute the abnormal behavior recognition method shown in some embodiments of this specification.
[0102] In some embodiments, obtaining dynamic features of the target detection object based on the captured image of the detection area includes:
[0103] Based on the collected image of the detection area, target information in the collected image is obtained;
[0104] Based on the target information, the dynamic characteristics of the target detection object are obtained.
[0105] Target information generally refers to various features related to the target detection object in the captured image. For example, target information may include the location information of the target detection object, the location information of relevant parts of the target detection object, etc.
[0106] In some embodiments, the target information includes detection boxes and / or key points.
[0107] A detection frame is a recognized image region that is used to define a specific target object in the captured image. The target object can include the face, hand, or other objects of the target detection subject.
[0108] Key points refer to key locations of the target object in the captured image. For example, key points can include key points on the hand of the target detection object, key points on the face of the target detection object, etc.
[0109] In some embodiments, target objects in an image (such as the face, hands, steering wheel, mobile phone, etc. of the target detection subject) can be located and identified based on a target detection algorithm (such as YOLO, Faster R-CNN, etc.) to generate multiple detection frames.
[0110] In some embodiments, the target detection algorithm can output relevant information of the detection frame, for example, the position of the detection frame in the captured image, which can be represented by a set of coordinates, such as the coordinates of the upper left corner and the lower right corner of the detection frame, or using the coordinates of the center point of the detection frame and the width and height of the detection frame. The size of the detection frame, i.e., the width and height, reflects the actual size of the target object in the captured image. For example, the category label associated with each detection frame indicates the category to which the target object contained in the detection frame belongs (for example, "hand", "face", "mobile phone", etc.). For example, the confidence score is a value between 0 and 1, which indicates the degree of confidence in the detection result. A high score indicates that the detection result is very likely to be correct, while a low score indicates that the detection result is less likely to be correct.
[0111] In some embodiments, key point information within the detection frame, such as the positions of facial feature points such as eyes, nose, mouth, and eyebrows, can be extracted through a key point detection algorithm (such as Dlib, MMPose, etc.).
[0112] Step 320: Identify abnormal behavior based on the dynamic characteristics of the target detection object.
[0113] Dynamic features refer to features extracted from multiple frames of images, which reflect the behavioral changes of the target detection object over a period of time.
[0114] In some embodiments, the behavior recognition result of the target detection object can be determined based on the dynamic features.
[0115] In some embodiments, dynamic features can be obtained in a variety of ways based on the key points of multiple frames of captured images. For example, the displacement speed of the key points of the target detection object between consecutive frames can be calculated to determine whether the target detection object's action is fast (such as quickly lowering the head, waving the hand quickly, etc.), or the speed change rate of the key points can be calculated to determine the acceleration of the action; or the motion trajectory of the key points in multiple frames can be determined to determine the behavior pattern of the target detection object (such as whether the head is frequently shaken left and right, whether the hand is frequently stretched out, etc.), or the behavior pattern of the target detection object can be analyzed based on the key point information of the previous and subsequent frames of the current frame capture image, such as whether the head is lowered for a long time, whether the mobile phone is frequently used, etc.
[0116] Through dynamic features, abnormal behaviors of target detection objects can be identified and warned more accurately, improving driving safety.
[0117] In some embodiments, the detection frame includes a detection frame of at least one of a face, a hand, and an abnormal object related to abnormal behavior of the target detection object.
[0118] Abnormal objects related to the abnormal behavior of the target detection object may include water bottles, cigarettes, and mobile phones.
[0119] In some embodiments, based on historical data or prior knowledge, abnormal behaviors of the target detection object can be determined, including talking (holding a mobile phone), speaking (holding a mobile phone), smoking, and drinking water. Based on the abnormal behaviors, the target objects of the DMS target detection algorithm are determined, including but not limited to the face, hands, water bottles, cigarettes, and mobile phones of the target detection object, and the target detection algorithm is used to obtain multiple detection frames in the captured image.
[0120] In some embodiments, image data can be collected, cleaned, and annotated based on the abnormal behavior of the target detection object. For example, a target detection object abnormal behavior set is established:
[0121] driver_dangeous_behavior={b1,b2...,b n};
[0122] Where n is the abnormal behavior category of the target detection object, which should include but not be limited to smoking, making phone calls or sending voice messages, and drinking water;
[0123] Establish a set of abnormal behavior scenarios for target detection objects:
[0124] driver_dangerous_behavior_sense={s1,s2,…,s m};
[0125] Where m is the category of the scene in which the abnormal behavior of the target detection object occurs, including but not limited to: daytime scene, nighttime scene, dusk / early morning scene, high-speed vehicle movement scene, and low-speed vehicle movement scene;
[0126] Combine the target detection object abnormal behavior set with the target detection object abnormal behavior scene set to obtain the data collection requirement table:
[0127] driver_dangerous_behavior_data_requirement={r1, r2,…, r l};
[0128] Among them, l is the number of abnormal behavior collection requirements of the target detection object. Data collection needs to be performed according to the data collection requirement table to obtain a preliminary data set; the obtained preliminary data set is cleaned and converted to obtain the final target detection object abnormal behavior detection data set.
[0129] Exemplarily, the elements in the target detection object abnormal behavior set and the target detection object abnormal behavior scene set can be arranged and combined to finally obtain the collected data collection requirement table. For example, the data collection requirement table may include: making phone calls during the day, sending voice messages during the day, playing with mobile phones during the day, no abnormal behavior during the day, making phone calls at night, playing with mobile phones at night, sending voice messages at night, no abnormal behavior at night, etc.
[0130] In some embodiments, the target detection object abnormal behavior detection dataset can be input into a pre-trained target detection algorithm and a key point detection algorithm to obtain the key points and detection boxes of the target object related to the abnormal behavior of the target detection object.
[0131] In some embodiments, inaccurate parts of the detection frame or key pre-labeled labels can be manually modified to obtain the labeled labels of the input data; for example, each frame of the captured image is labeled, and its labeling categories include but are not limited to smoking, making phone calls / chatting, eating, non-risky behaviors, unlabeled / invalid data, etc., where the data of the first n frames and the last n frames of the captured image of each frame should be labeled as unlabeled / invalid data. Because the DMS algorithm combines the data of the first n frames and the last n frames, at least 2n+1 frames of image video are required to output the results of the n+1 frame of the captured image. For the first n frames and the last n frames in each image video, due to the lack of sufficient data before and after the frames, the DMS algorithm cannot perform a complete analysis and judgment.
[0132] In some embodiments, a set of detection frames related to abnormal behaviors related to the target detection object can be obtained based on the detection frames of the target object (specific parts of the target detection object, abnormal objects, etc.) related to abnormal behaviors in the current frame captured image detected by the target detection algorithm, including but not limited to the face detection frame, hand detection frame, mobile phone detection frame, cigarette detection frame, water bottle detection frame, or water cup detection frame of the target detection object:
[0133] detect_bbox_set={face_bbox,left_hand_bbox,right_hand_bbox,cigarette_bbox,phone_bbox,bottle_bbox};
[0134] Among them, detect_bbox_set is the set of detection frames related to the abnormal behavior of the target detection object, face_bbox is the face detection frame of the target detection object, left_hand_bbox and right_hand_bbox are the left hand detection frame and right hand detection frame of the target detection object, cigarette_bbox is the cigarette detection frame, phone_bbox is the mobile phone detection frame, and bottle_bbox is the detection frame related to drinking water, such as water cups and water bottles. Each detection frame can include the following information:
[0135] bbox = {x, y, w, h};
[0136] Among them, x, y, w and h are the horizontal and vertical coordinates of the center point of the detection frame and the width and height information of the detection frame respectively.
[0137] In some embodiments of this specification, by fusing the dynamic features and static features of the target detection object, the final behavior recognition result of the target detection object is made more reliable, reducing the misclassification problem of dangerous behaviors caused by target object occlusion and different angles.
[0138] In some embodiments, obtaining dynamic features of a target detection object based on target information includes:
[0139] Determining target key points in the captured image that correspond to target parts of the target detection subject's body;
[0140] Based on the target key point information in each frame of the image, the motion information of the target part is obtained to determine the dynamic characteristics.
[0141] Target key points refer to the feature points of the target part of the target detection object. The target key points may include feature points corresponding to the face (such as eyes, nose, mouth, etc.), hands (such as wrists, finger joints, etc.) of the target detection object.
[0142] Target key point information refers to the data related to one or more feature points of a specific part of the target detection object extracted from the captured image. Target key points can include specific locations on the target detection object's face, hands, and other target parts.
[0143] In some embodiments, the target key point information includes position information of the target key point.
[0144] The location information of the target key point refers to the coordinates of the target key point in the image.
[0145] In some embodiments, each frame of the captured image may be processed to extract target key point information of the target portion of the target detection object. In some embodiments, the target key point information of the target portion of the target detection object may be obtained through other systems (e.g., a driver monitoring system).
[0146] Motion information refers to the motion change information of each of the multiple target key points of the target part between consecutive frames. For example, the motion information may include the speed, acceleration, and motion trajectory of the target key point.
[0147] In some embodiments, for each target key point of the target part, the velocity of the target key point is obtained based on the displacement and time interval of the target key point between multiple frames of acquired images; the velocity change rate of the target key point is calculated to obtain the acceleration of the key point; and the motion path of the key point is obtained based on the position change of the target key point in multiple frames of acquired images.
[0148] In some embodiments of this specification, by extracting key point information from each frame of the captured image and calculating motion information based on the key point information, dynamic features can be accurately determined, which helps to effectively identify dangerous behaviors of the target detection object and provide strong protection for driving safety.
[0149] In some embodiments, obtaining the motion information of the target part based on the target key point information in each frame of the acquired image includes:
[0150] Based on the position change information of the target key points in multiple frames of collected images, the motion information of the target part is obtained.
[0151] Position change information refers to the position difference of the target key point between consecutive frame images. For example, the position change information may include the distance between the target key point and the consecutive frame images.
[0152] In some embodiments, for each target key point of the target part, the current frame acquisition image is matched with the target key point in the previous frame acquisition image to achieve tracking of each target key point in multiple consecutive frame acquisition images, and the displacement vector of the target key point between every two frame acquisition images is calculated, and the speed of the target key point is obtained by dividing the displacement vector by the frame time interval.
[0153] In some embodiments, for each target key point of the target part, the Lucas-Kanade optical flow or other optical flow algorithms may be used to estimate the movement direction and speed of the target key point.
[0154] In some embodiments of this specification, the motion state of the target detection object can be comprehensively and accurately analyzed through position change information.
[0155] In some embodiments, obtaining the motion information of the target part based on the position change information of the key points in the multiple frames of acquired images includes:
[0156] A plurality of groups of key frame images with different time intervals are acquired, and first motion information and second motion information are obtained based on position change information of target key points in the plurality of groups of key frame images.
[0157] In some embodiments, frame images with different time intervals can be selected in the image video, for example, a frame of captured image can be selected every 1 second, 2 seconds or longer time intervals to obtain a set of key frame images, so as to effectively reduce the amount of calculation and capture the motion changes of the target key points over a longer period of time.
[0158] In some embodiments, a key frame image can be dynamically selected based on the motion state (e.g., speed, acceleration) of a target key point in an image video. For example, when the speed or acceleration of a target key point exceeds a certain threshold, the frame is selected as the key frame image.
[0159] Multiple sets of key frame images refer to a set of key frames selected in a video sequence according to a specific time interval or other strategy.
[0160] The first motion information and the second motion information refer to the motion information of the target key point in different time intervals. For example, the first motion information refers to the motion information of the target key point in a shorter time interval, and the second motion information refers to the motion information of the target key point in a longer time interval.
[0161] In some embodiments, obtaining the motion information of the target part based on the position change information of the target key points in the multiple frames of acquired images includes:
[0162] First motion information is obtained based on position change information of the target key point in a group of key frame images within a first time interval.
[0163] The first time interval refers to a shorter time interval and can be determined based on experiments or experience.
[0164] The first motion information may consist of rapid motion speeds of multiple target key points of the same target part.
[0165] Extraction of primary motion information primarily utilizes a shorter inter-frame interval to extract rapid motion information from the target's key points, capturing rapid or transient changes in the target's movements. This information helps the detection model understand the rapid movement and momentary posture changes associated with abnormal behavior, facilitating the extraction of transient abnormal behavior. By focusing on rapidly changing key points, the detection model can more sensitively capture subtle differences in movement, thereby improving the accuracy of identifying abnormal behavior and preventing even fleeting incidents of abnormal behavior.
[0166] In some embodiments, obtaining the motion information of the target part based on the position change information of the target key points in the multiple frames of acquired images further includes:
[0167] Second motion information is obtained based on position change information of the target key point in a group of key frame images at a second time interval, where the first time interval is shorter than the second time interval.
[0168] The second time interval refers to a longer time interval and can be determined based on experiments or experience.
[0169] The second motion information may consist of slow motion speeds of multiple target key points of the same target part.
[0170] Compared with the first motion information, the second motion information uses a relatively longer inter-frame interval to extract the slow motion information of the target detection object, which mainly reflects the overall posture and position of the target part related to the abnormal behavior of the target detection object, rather than the specific action details. It helps the detection model understand the overall structure and continuity characteristics of the gesture, and helps to extract long-term action change information. The slow motion information can provide a stable reference point, helping the detection model maintain its understanding of the abnormal behavior of the target detection object under complex background or lighting conditions, and reduce misjudgments caused by short-term fluctuations.
[0171] Figure 4 This is an exemplary schematic diagram of determining motion information according to some embodiments of this specification.
[0172] In some embodiments, as Figure 4 As shown, the key point detection algorithm is used to obtain the target key point information of the face, left hand and right hand of the target detection object of the current frame acquisition image (denoted as the i-th frame), the previous frame acquisition image (denoted as the in-th frame), the previous frame acquisition image (denoted as the im-th frame), the subsequent frame acquisition image (denoted as the i+m-th frame) and the subsequent frame acquisition image (denoted as the i+n-th frame), such as:
[0173] key_point i={face_key_point int i ,left_hand_key_pointi,right_hand_key_p0int i};
[0174] Among them, key_point i The target key point information of the i-th frame, the target key point information of the in-th frame, im-th frame, i+m frame and i+n frame is key_point in the same way i-n 、key_point i-m 、key_point i+m and key_point i+n , and n>m>0, face_key_point i The target key point information of the target detection object's face in the current frame image is collected, which may include: the three-dimensional spatial position information of the target key point of the target detection object's face and whether it is visible. The number of target key points depends on the definition of the target detection object's facial key points by the key point detection algorithm, left_hand_key_point i and right_hand_key_point i They are the target key point information of the left hand and the target key point information of the right hand of the target detection object in the current frame acquisition image, and the information they contain is similar to the target key point information of the face of the target detection object, such as:
[0175]
[0176] Among them, f_x i,j_f 、f_y i,j_f 、f_z i,j_f and f_v i,j_f are the horizontal coordinate information, vertical coordinate information, depth information and visibility information of the target key point on the face of the target detection object in the i-th frame acquisition image. If the current target key point is visible, the value of v is 1, and if it is invisible, the value of v is 0; the target key points on the left hand and the right hand are the same; if both target key points are invisible, the horizontal and vertical coordinates and depth information corresponding to the target key points are defined as 0;
[0177] In some embodiments, face_key_point i-n 、face_key_point i and face_key_point i+n Calculate the first motion information (fast motion information) of the face of the corresponding target detection object in the i-th frame captured image.
[0178]
[0179] Among them, ff_s i,j_ff,l is the fast motion speed of the j_ffth target key point on the face of the target detection object between the previous frame acquisition image and the i-th frame acquisition image, ff_s i,j_ff,r The fast motion speed of the j_ffth target key point on the face of the target detection object between the subsequent frame acquisition image and the i-th frame acquisition image is calculated as follows:
[0180]
[0181] ∑f(v i )!=0,2f(v i-n )! = 0;
[0182] Among them, ff_s i,j_ff,l is the fast motion speed of the j_ffth target key point on the face of the target detection object between the i-th frame acquisition image and the in-th frame acquisition image, f_x i,j_ff 、f_y i,j_ff and f_z i,j_ff The horizontal coordinate information, vertical coordinate information and depth information of the j_ffth target key point on the face of the target detection object in the i-th frame acquisition image are as follows. Similarly, f_x i-n,j_ff 、f_y i-n,j_ff and f_z i-n,j_ff The information corresponding to the j_ffth target key point on the face of the image is collected for the inth frame, ff_s i,j_ff,r The fast motion speed of the j_ffth target key point on the face of the target detection object between the i-th frame acquisition image and the i+n-th frame acquisition image, which is calculated in the same way as ff_s i,j_ff,l Similar.
[0183] In some embodiments, fs_s i,j_fs,l and fs_s i,j_fs,r They are the slow motion speed of the j_fsth target key point on the face of the target detection object between the previous frame acquisition image and the i-th frame acquisition image, and the slow motion speed of the j_fsth target key point on the face of the target detection object between the subsequent frame acquisition image and the i-th frame acquisition image. The calculation method is similar to the fast motion speed. The second motion information of the face of the target detection object in the i-th frame acquisition image is as follows:
[0184]
[0185] In some embodiments of the present specification, multiple sets of key frame images can be used to reduce misjudgments caused by noise or abnormal data in a single frame image; by extracting motion information from key frame images at different time intervals, it can better adapt to complex driving environments and lighting conditions and improve the robustness of dangerous behavior recognition.
[0186] In some embodiments, the method further comprises:
[0187] The dynamic features are determined based on the motion information of the target part and the distance information of the key points of the target part.
[0188] Each element in the key point distance information represents the distance between two target key points in the same target part. The key point distance information can represent the relative position relationship between different target key points in the target part.
[0189] Exemplarily, based on the target key point information of the face of the target detection object in the i-th frame captured image, the key point distance information of the face of the target detection object in the i-th frame captured image is calculated as follows:
[0190]
[0191] In some embodiments, key point distance information is obtained by:
[0192] Based on the distance between every two target key points of the target part, key point distance information of the target part is determined.
[0193] Among them, d j_c,k_c is the distance between the j_cth target key point and the k_cth target key point. The distance calculation methods include but are not limited to Manhattan distance, Euclidean distance, and Chebyshev distance. Taking Euclidean distance as an example, its calculation method is as follows:
[0194]
[0195] Among them, f_x i,j_c and f_x i,k_c The j_cth target key point and the k_cth target key point of the face of the target detection object in the i-th frame acquisition image, d j_c,k_c =d k_c,j_c And when j_c==k_c, d j_c,k_c =0; therefore face_distance i For a matrix whose main diagonal is all zero and is symmetric about the main diagonal, assign all values of its upper triangular area to 0 and convert it into a one-dimensional vector.
[0196] In some embodiments, the first motion information, second motion information and key point distance information of the left hand and the right hand are calculated in a similar manner to that of the face of the target detection object. When the hands are completely invisible, the corresponding first motion information, second motion information and key point distance information are all assigned a value of -1.
[0197] Keypoint distance information refers to the relative distances between keypoints on a specific target within the same frame. This information is primarily used to describe the target object's shape, posture invariance, and spatial relationships. For example, keypoint distance information can accurately describe the shape of the target object at a specific moment, helping to distinguish similar behaviors from those involved in the target's abnormal behavior. Furthermore, even if the target object associated with the target's abnormal behavior rotates or translates, the relative distances between keypoints remain unchanged. Therefore, keypoint distance information helps identify abnormal behaviors with unchanged postures. Furthermore, by analyzing the distances between keypoints, the detection model can better understand the internal spatial relationships of the target object associated with the target's abnormal behavior, improving its ability to detect details related to the target's abnormal behavior.
[0198] In some embodiments of this specification, the shape of the target part can be analyzed through key point distance information, and dynamic features can be extracted to more accurately identify and warn of dangerous behaviors and ensure driving safety.
[0199] In some embodiments, the method further comprises:
[0200] Feature extraction is performed on the motion information of the target part and the distance information of the key points of the target part respectively to obtain the motion feature of the target part and the distance feature of the key points of the target part;
[0201] The dynamic features are determined based on the motion features and key point distance features of each target part.
[0202] In some embodiments, based on deep learning model extraction, the motion information of the target part and the key point distance information of the target part can be processed separately to obtain corresponding feature vectors as the motion features of the target part and the key point distance features of the target part.
[0203] In some embodiments of this specification, by combining the motion information (speed, acceleration) of the target part with the distance information (relative position changes of key points) of the target part, a more comprehensive analysis of the target detection object's behavior can be achieved. For example, speed alone may not be able to fully identify certain complex behaviors, but combining relative position changes can improve recognition accuracy. By integrating multi-dimensional features, the target detection object's behavior can be more accurately judged, reducing misjudgments caused by a single feature.
[0204] In some embodiments, the method further comprises:
[0205] Based on the collected images of the detection area, the static features of the target detection object are obtained;
[0206] Identify abnormal behaviors based on the dynamic and static features of the target detection object.
[0207] Static features refer to features extracted from a single frame of captured image, which reflect the state of the target detection object at a certain moment.
[0208] In some embodiments, static features can be obtained in a variety of ways based on the detection frame in the captured image. For example, features related to the target detection object can be extracted within the detection frame: using a key point detection algorithm (such as Dlib, MTCNN, etc.) to extract facial expression features of the target detection object, such as the degree of eye opening and closing, the opening and closing state of the mouth, etc., or extracting the position information of the key points of the target detection object's hand within the corresponding detection frame to determine whether the hand is on the steering wheel; or using a posture estimation algorithm to extract the posture of the target detection object's head.
[0209] In some embodiments, the processor may also process the dynamic features and static features of the target detection object through a detection model to determine the behavior recognition result of the target detection object. The detection model may be a machine learning model such as a trained neural network. The input of the detection model may include the dynamic features and static features of the target detection object, and the output may include the behavior recognition result of the target detection object. The detection model may be trained based on a large number of labeled training samples through various feasible methods. For example, parameter updates may be performed based on a gradient descent method. The training samples may include the dynamic features and static features of the sample acquisition image, which may be obtained based on historical data. The labels may be the actual behavior recognition results of the target detection object corresponding to the sample acquisition image, which may be obtained by manual or automatic annotation.
[0210] In some embodiments, abnormal behavior of the target detection object can be identified directly based on the behavior recognition result.
[0211] In some embodiments, the static features include image features of a target area of the target detection object in the captured image, and the method further includes:
[0212] Based on multiple detection frames in the captured image, determine the detection frame closest to the target detection frame;
[0213] Based on a detection frame closest to the target detection frame, image features of a target area of the target detection object are determined, wherein the target area includes at least one target part of the target detection object.
[0214] The target region refers to a specific area in the captured image that is related to the target detection object. For example, the target region may include the entire body region of the target detection object, the facial region of the target detection object, or a specific part of the target detection object.
[0215] In some embodiments, the target area is one of a face area, a hand area, and a whole body area of the target detection object.
[0216] In some embodiments, the target area is a facial area of the target detection object, and the target detection frame is a facial detection frame.
[0217] In some embodiments, determining a detection frame closest to a target detection frame based on multiple detection frames in a captured image includes:
[0218] Based on the distance between the target detection frame and other detection frames and a preset distance threshold, the detection frame closest to the target detection frame is determined.
[0219] The preset distance threshold is a preset critical value of the distance between detection frames. The preset distance threshold can be a system default value, a manually preset value, etc.
[0220] Exemplarily, the obtained multiple detection frames are fused. For example, detection frames with no detection information other than the target detection frame (e.g., the face detection frame face_bbox in the captured image) are removed, and the distance between the remaining detection frames and the target detection frame is determined, and detection frames that are far from the target detection frame are filtered out. Distance metrics between detection frames include, but are not limited to, Manhattan distance, Euclidean distance, and Intersection over Union (IoU) distance.
[0221] For example, the target detection frame is a face detection frame, and the distance between each detection frame and the face detection frame is as follows:
[0222] distance_s et
[0223] ={left_hand_dis,right_hand_dis,cigarette_dis,phone_dis,bottle_dis};
[0224] The distance calculation method takes Euclidean distance as an example. The distance between detection frames is as follows:
[0225]
[0226] Where dis(face_bbox, other_bbox) is the distance between the face detection box of the target detection object and other detection boxes. If the other detection boxes are empty, the distance value is -1. Otherwise, the distance between the detection boxes is the Euclidean distance between the center points of the detection boxes.
[0227] In some embodiments, the distances between each detection frame and the target detection object's face detection frame are obtained and filtered based on a preset distance threshold to obtain a combination of the detection frames closest to the face detection frame, i.e., a set of target detection frames related to the target detection object's abnormal behavior:
[0228] dangerous_behavior_a bout_bbox_set={bbox1, bbox2,...,bbox k};
[0229] In some embodiments of this specification, by selecting a detection frame closest to the face detection frame, the target area can be located more accurately and misjudgment can be reduced.
[0230] In some embodiments, determining image features of a target area of a target detection object based on a detection frame closest to the target detection frame includes:
[0231] Merge the detection frame closest to the target detection frame with the target detection frame to obtain the circumscribed rectangular frame;
[0232] Feature extraction is performed on the image area corresponding to the circumscribed rectangular frame to obtain the image features of the target area of the target detection object.
[0233] Figure 5 This is an exemplary schematic diagram of image features of a target area of a target detection object according to some embodiments of this specification.
[0234] For example, the target area is the face area, such as Figure 5 As shown in the figure, the minimum vertical bounding box between all detection boxes in the target detection box set and the face detection box is obtained, and it is expanded by a certain proportion according to its width and height information, which is recorded as driver_dangerous_about_max_bbox, which includes the following information:
[0235] driver_dangerous_behavior_about_bbox={x,y,w,h};
[0236] Among them, x, y, w, and h are the horizontal coordinate information / vertical coordinate information and width and height information of the center point of driver_dangerous_behavior_about_bbox respectively.
[0237] In some embodiments, the captured image of the current frame is cropped based on the obtained driver_dangerous_about_max_bbox to obtain an image area corresponding to the circumscribed rectangular frame.
[0238] In some embodiments, the image area corresponding to the circumscribed rectangular box of a pre-trained convolutional neural network (such as VGG, ResNet, etc.) can be used for feature extraction, and the output before the last fully connected layer is used as the feature vector of the image area corresponding to the circumscribed rectangular box to obtain the image features of the target area of the target detection object.
[0239] In some embodiments, the size of the image area corresponding to the circumscribed rectangular frame can be adjusted to an image of a predefined size; the image of the predefined size is input into a trained convolutional neural network (CNN) to obtain image features of the target area of the target detection object in the current frame capture image.
[0240] In some embodiments of this specification, by merging detection frames, more comprehensive coverage of the target area of the target detection object can be achieved, reducing missed detections due to partial occlusion or misjudgment. For example, if the face detection frame and the hand detection frame are close together, the merged bounding rectangle can include information about both the face and the hand, avoiding misjudgments caused by the hand occluding the face. The merged bounding rectangle can also more accurately reflect the facial and hand positions of the target detection object, thereby improving detection accuracy.
[0241] In some embodiments, the static features include relative position features,
[0242] The method also includes:
[0243] Based on multiple detection frames in the captured image, the distance between each two detection frames is calculated to determine the detection frame distance matrix;
[0244] Feature extraction is performed on the detection frame distance matrix to obtain relative position features.
[0245] In some embodiments, for each detection frame in the captured image, the distance between each detection frame and the other detection frames is calculated. For example, the distance between the center points of two detection frames can be calculated based on the Euclidean distance formula.
[0246] Figure 6 is an exemplary flow chart of relative position features according to some embodiments of this specification.
[0247] In some embodiments, as Figure 6As shown, the detection frame set related to the abnormal behavior of the target detection object of the current frame acquisition image is obtained, and the corresponding detection frame matrix is constructed. The detection frame matrix form is as follows:
[0248]
[0249] Among them, fb, lhb, rhb, cb, pb and bb represent the target detection object face detection frame, left hand detection frame, right hand detection frame, cigarette detection frame, mobile phone detection frame and water bottle detection frame respectively.
[0250] The detection frame matrix is a matrix formed by pairing different detection frames.
[0251] In some embodiments, based on the obtained detection box matrix, combined with various distance calculation formulas, the distance between different detection boxes is obtained, that is, the detection box distance information bbox_iou_matrix. Among them, various distance calculation formulas include but are not limited to: IoU distance, GIoU (Generalized Intersection over Union) distance, DIoU (Distance Intersection over Union) distance, CIoU (Complete Intersection over Union) distance and EIoU (Enhanced Intersection over Union) distance. Taking IoU distance as an example, bbox_iou_matrix is calculated as follows:
[0252]
[0253] In some embodiments, the elements on the main diagonal of the detection box distance information bbox_iou_matrix are set to 1, and IoU(bbox1, bbox2) = IoU(bbox2, bbox1), that is, the detection box distance information bbox_iou_matrix is a matrix symmetric about the main diagonal. The values in the lower triangular area of the main diagonal of the detection box distance information are converted to a one-dimensional vector to obtain bbox_iou_vector; the obtained bbox_iou_vector is input into a pre-trained FCN (Fully Convolutional Networks) to obtain relative position features.
[0254] By determining relative position features, the detection model learns the importance of different detection box positions, allowing it to focus more on areas critical to the classification task when generating feature maps. This helps reduce the impact of background noise and highlight key parts of the target object, thereby improving classification accuracy. It also helps the detection model better adapt to varying image content and layout variations. When faced with images with similar backgrounds but different target objects, the detection model can still accurately capture the main features, improving its generalization ability. In complex scenarios, target objects may appear at different scales. Incorporating intersection-over-union (IoU) information (i.e., the relative distance between detection boxes) allows the attention weights to be dynamically adjusted based on the feature maps at each scale, ensuring that the detection model can effectively extract features at multiple scales. This is particularly useful for classification tasks involving objects of varying sizes. Since the incorporation of IoU information allows the detection model to focus on more important areas, it can reduce unnecessary computation, particularly in resource-constrained environments such as mobile devices or embedded systems.
[0255] In some embodiments of this specification, by calculating the distance between each two detection frames, the relative position changes between the target detection subject's face, hands, head, and other parts can be captured, helping to identify complex behavioral patterns (such as distracted driving and fatigue driving). Relative position features can provide more comprehensive contextual information and reduce misjudgments caused by a single detection frame. For example, if a hand detection frame is close to a face detection frame, it may indicate that the target detection subject is touching the face.
[0256] In some embodiments, identifying abnormal behavior includes:
[0257] Feature fusion is performed based on static features and dynamic features to determine the behavior recognition results of the target detection object.
[0258] In some embodiments, weighted operations may be performed on features such as static features and dynamic features; the weighted features may be concatenated; and the concatenated features may be input into a fully connected layer for classification to obtain a final behavior recognition result.
[0259] The behavior recognition result integrates the fast and slow motion information of the target detection object, the relative position information of each detection frame, and the image features of the current frame, making the final motion information more reliable and minimizing the misclassification of abnormal behaviors caused by occlusion and different angles.
[0260] In some embodiments, feature fusion is performed based on static features and dynamic features to determine the behavior recognition result of the target detection object, including:
[0261] Based on the first weight and the second weight, the dynamic features and the static features are weightedly fused to obtain fused features;
[0262] Perform feature extraction on the fused features to determine the behavior recognition results of the target detection object.
[0263] The first weight and the second weight may be determined based on experiments or experience.
[0264] For example, the dynamic features and static features are converted into one-dimensional feature vectors. For example, the image features, relative position features, motion features of the target area of the target detection object, and key point distance features of the face and hands of the target detection object are converted into one-dimensional vectors; the one-dimensional feature vectors are weighted and fused to form fused features, and the weighted fusion method is as follows:
[0265] multi_modal_data=α*image_feature+β*IoU_feature+γ*key_point_feature;
[0266] Among them, multi_modal_data is the fusion feature, α, β and γ are the image features, relative position features and dynamic features of the target area of the target detection object respectively; among them, the dynamic features include the motion features and key point distance features of the target parts such as the face and hands of the target detection object.
[0267] In some embodiments, the detection model is used to perform feature extraction and multimodal fusion on the fused features to obtain the final behavior recognition result of the target detection object.
[0268] It should be noted that the dimensions of the image features, relative position features, motion features of the face and hands of the target detection object, and key point distance features of the target area of the target detection object can be the same or different.
[0269] Figure 7 is an exemplary schematic diagram of weighted fusion according to some embodiments of this specification.
[0270] like Figure 7As shown, in some embodiments, the image features, relative position features, motion features of the target area of the target detection object, and key point distance features of the target detection object are multiplied by different weights, and the weighted image features, relative position features, motion features of the target area of the target detection object, and key point distance features are concatenated (e.g., concatenated based on dimension, [1, 2] + [2, 3, 4] + [9] = [1, 2, 2, 3, 4, 9]) to obtain fused features. Specific provisions are not made here, and subsequent detection models can achieve the final driving behavior classification task for input one-dimensional vectors.
[0271] For the image features of the target area of the target detection object in the current frame capture image, the relative position features of the target detection object, the motion features of the target parts such as the face and hands of the target detection object, and the key point distance features of the target parts, it contains image information, target position information and key point information, the feature types are richer, and the classification results are more accurate.
[0272] In some embodiments of this specification, dynamic features (such as motion information) and static features (such as image features within the detection frame) each provide different information. Through weighted fusion, dynamic information and static information can be fully utilized to improve the accuracy of driving behavior classification; dynamic features can capture instantaneous behavior, while static features can provide stable contextual information. Combining the two can reduce misjudgments caused by a single feature.
[0273] It should be noted that the above description of the relevant processes is for illustration and purpose only and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the processes under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.
[0274] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0275] Figure 8 This is a structural diagram of an electronic device according to some embodiments of this specification.
[0276] The embodiment of the present application further provides an electronic device 800, which may include one or more processors 801 of processing cores, one or more computer-readable storage media memories 802, a power supply 803, an input unit 804, and other components. Those skilled in the art will appreciate that Figure 8 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0277] The processor 801 is the center of the abnormal behavior identification system. It uses various interfaces and lines to connect the various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 802, and calling the data stored in the memory 802, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. It is understood that the processor 801 transmits signals with the controller. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application programs, etc., and the modem processor mainly processes wireless communications. It is understood that the above-mentioned modem processor may not be integrated into the processor 801.
[0278] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.
[0279] In some embodiments of the present application, the abnormal behavior identification device can be implemented in the form of a computer program. The computer program can be used in Figure 8 The electronic device is operated on the device shown. The memory of the electronic device may store the various program modules that constitute the abnormal behavior identification device. The computer program composed of the various program modules enables the processor to execute the steps of the abnormal behavior identification method of each embodiment of the present application described in this specification.
[0280] The electronic device includes a processor, memory, and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with external electronic devices via a network connection. When executed by the processor, the computer program implements a method for identifying abnormal behavior.
[0281] The electronic device also includes a power supply 803 for supplying power to various components. Preferably, the power supply 803 can be logically connected to the processor 801 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 803 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0282] The electronic device may further include an input unit 804, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0283] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the electronic device loads the executable files corresponding to one or more application processes into the memory 802 according to computer instructions, and the processor 801 runs the application stored in the memory 802, thereby implementing various functions, such as the abnormal behavior identification method of each embodiment of the present application described in this specification.
[0284] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0285] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to implement as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments and will not be repeated here.
[0286] It should be noted that Figure 8 This is only one implementation of the electronic device 800 provided in the embodiment of the present application. In actual applications, the electronic device 800 may also include more or fewer components, which is not limited here.
[0287] It should be understood that the various schemes of the embodiments of the present application can be reasonably combined and used, and the explanations or descriptions of the various terms appearing in the embodiments can be referenced or explained with each other in the various embodiments, without limitation to this.
[0288] It should also be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0289] Based on the above embodiments and the same concept, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer, the computer executes the method provided in the above embodiments.
[0290] Based on the above embodiments and the same concept, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer executes the method provided in the above embodiments.
[0291] It is understandable that, in order to implement the functions of any of the above-mentioned embodiments, the electronic device (such as a vehicle-mounted terminal) includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0292] Figure 9 is an exemplary schematic diagram of a vehicle according to some embodiments of the present specification.
[0293] like Figure 9 As shown, the embodiment of the present application further provides a vehicle, which includes the electronic device described in any embodiment. The vehicle can be a fuel vehicle, a plug-in hybrid vehicle, or a new energy vehicle, etc., which is not specifically limited in this specification.
[0294] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0295] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.
[0296] The above are only preferred embodiments of the present application and do not constitute any form of limitation to the present application. Although the descriptions of each embodiment in the embodiments of the present application have different focuses, for parts that are not described in detail in a certain embodiment, please refer to the relevant embodiments of other embodiments. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A method for identifying abnormal behavior, characterized in that: The method comprises: Based on the collected images of the detection area, the dynamic characteristics of the target detection object are obtained; the detection area is the area where the abnormal behavior is observed; Abnormal behavior is identified based on the dynamic characteristics of the target detection object.
2. The method according to claim 1, characterized in that The acquisition of dynamic features of the target detection object based on the collected image of the detection area includes: Based on the collected image of the detection area, obtaining target information in the collected image; Based on the target information, dynamic features of the target detection object are obtained.
3. The method according to claim 2, characterized in that The target information includes a detection box and / or key points.
4. The method according to claim 2, characterized in that The obtaining of dynamic features of the target detection object based on the target information includes: Based on the target information, determining a target key point in the acquired image corresponding to a target part of the target detection object; Based on the target key point information in each frame of the collected image, the motion information of the target part is obtained to determine the dynamic feature.
5. The method according to claim 4, characterized in that The target key point information includes position information of the target key point.
6. The method according to claim 4, characterized in that The obtaining of the motion information of the target part based on the target key point information in each frame of the acquired image includes: Based on the position change information of the target key point in the multiple frames of the acquired image, the motion information of the target part is obtained.
7. The method according to claim 6, characterized in that The obtaining of the motion information of the target part based on the position change information of the target key point in the multiple frames of the acquired image includes: First motion information is obtained based on position change information of the target key point in a group of key frame images within a first time interval.
8. The method according to claim 7, characterized in that The obtaining of the motion information of the target part based on the position change information of the target key point in the multiple frames of the acquired image further includes: Second motion information is obtained based on position change information of the target key point in a group of key frame images at a second time interval, where the first time interval is smaller than the second time interval.
9. The method according to claim 4, characterized in that The method further comprises: The dynamic feature is determined based on the motion information of the target part and the key point distance information of the target part.
10. The method according to claim 9, characterized in that The key point distance information is obtained by: Based on the distance between every two target key points of the target part, key point distance information of the target part is determined.
11. The method according to claim 1, wherein The method further comprises: Based on the collected images of the detection area, the static features of the target detection object are obtained; The abnormal behavior is identified based on the dynamic features and the static features of the target detection object.
12. The method according to claim 11, characterized in that The static features include image features of the target area of the target detection object in the collected image, The method further comprises: Determine, based on multiple detection frames in the captured image, a detection frame closest to the target detection frame; Based on the detection frame closest to the target detection frame, image features of a target area of the target detection object are determined, wherein the target area includes at least one target part of the target detection object.
13. The method according to claim 12, characterized in that The determining, based on the multiple detection frames in the acquired image, the detection frame closest to the target detection frame includes: Based on the distance between the target detection frame and other detection frames and a preset distance threshold, the detection frame closest to the target detection frame is determined.
14. The method according to claim 12, characterized in that The determining, based on the detection frame closest to the target detection frame, the image features of the target area of the target detection object includes: Merge the detection frame closest to the target detection frame with the target detection frame to obtain a circumscribed rectangular frame; Feature extraction is performed on the image area corresponding to the circumscribed rectangular frame to obtain image features of the target area of the target detection object.
15. The method according to claim 12, characterized in that The target area is one of a face area, a hand area, and a whole body area of the target detection object.
16. The method according to claim 15, characterized in that The target area is a face area of the target detection object, and the target detection frame is a face detection frame.
17. The method according to claim 11, characterized in that The static features include relative position features, The method further comprises: Based on the plurality of detection frames in the captured image, calculating the distance between each two detection frames to determine a detection frame distance matrix; Feature extraction is performed on the detection frame distance matrix to obtain the relative position feature.
18. The method according to any one of claims 1 to 17, characterized in that The identifying of abnormal behavior includes: Feature fusion is performed based on static features and dynamic features to determine the behavior recognition result of the target detection object.
19. The method according to claim 18, characterized in that The step of performing feature fusion based on static features and dynamic features to determine the behavior recognition result of the target detection object includes: Based on the first weight and the second weight, the dynamic feature and the static feature are weightedly fused to obtain a fused feature; Feature extraction is performed on the fused features to determine a behavior recognition result of the target detection object.
20. The method according to any one of claims 1 to 19, characterized in that The detection frame includes a detection frame of at least one of a face, a hand, and an abnormal object related to abnormal behavior of the target detection object.
21. The method according to any one of claims 1 to 19, characterized in that The detection area is the driving area of the vehicle, and the target detection object is the driver.
22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a computer, the computer is caused to implement the abnormal behavior identification method according to any one of claims 1 to 21.
23. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, enable the computer to implement the abnormal behavior identification method according to any one of claims 1 to 21.
24. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the abnormal behavior identification method according to any one of claims 1 to 21.
25. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 24.