Driver Abnormal Behavior Detection Methods and Electronic Devices

By using a deep learning network to detect abnormal behavior from multiple driver images and comprehensively analyzing the abnormal behavior of the drivers, the problem of inaccurate detection results in existing technologies is solved, and higher detection accuracy is achieved.

CN115019288BActive Publication Date: 2025-10-28CNAUTOCHIPS SHANGHAI CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110246957.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-05
Publication Date
2025-10-28
Estimated Expiration
2041-03-05

AI Technical Summary

Technical Problem

Existing methods for detecting abnormal driver behavior are not very accurate.

Method used

A deep learning network is used to detect abnormal behavior in multiple driver images. The network detects abnormal behavior in each image and then analyzes the abnormal behavior of the driver based on the detection results of multiple images.

Benefits of technology

This improves the accuracy of detecting abnormal driver behavior, further enhancing the accuracy of detection results compared to single-image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019288B_ABST
    Figure CN115019288B_ABST
Patent Text Reader

Abstract

This application discloses a method and electronic device for detecting abnormal driver behavior. The method comprises: acquiring multiple first images, wherein the first images include at least one part of the driver; performing abnormal behavior detection on each first image using a detection network to obtain a detection result for each first image; and obtaining a driver abnormal behavior detection result based on the detection result for each first image. This method can improve the accuracy of the abnormal driver behavior detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular to a method and electronic device for detecting abnormal driver behavior. Background Technology

[0002] A driver's behavior directly impacts driving safety. Abnormal behavior while driving can significantly increase the likelihood of a traffic accident. For example, smoking or using a mobile phone while driving distracts the driver and increases the probability of an accident.

[0003] Therefore, it is essential to detect abnormal behavior in drivers to determine if any abnormal behavior is detected, so that drivers can be alerted and corrected if such behavior is found.

[0004] However, existing methods for detecting abnormal driver behavior do not yield accurate results. Summary of the Invention

[0005] This application provides a method and electronic device for detecting abnormal driver behavior, which can solve the problem that the detection results obtained by existing methods for detecting abnormal driver behavior are not accurate.

[0006] To address the aforementioned technical problems, this application provides a method for detecting abnormal driver behavior. The method includes: acquiring multiple first images, each containing at least one part of the driver's body; performing abnormal behavior detection on each first image using a detection network to obtain a detection result for each first image; and obtaining a result for detecting abnormal driver behavior based on the detection result of each first image.

[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a processor and a memory connected to the processor, wherein the memory stores program instructions; the processor is used to execute the program instructions stored in the memory to implement the above-mentioned method.

[0008] In this application, a detection network is used to detect abnormal behavior from multiple first images, and the abnormal behavior of the driver is comprehensively analyzed based on the detection results of the multiple first images. Since the detection network is a deep learning network, it can improve the accuracy of the detection results of the first images compared to traditional image processing methods, thereby improving the accuracy of the abnormal behavior detection results of the driver obtained based on the detection results of the first images. Furthermore, compared to analyzing abnormal driver behavior using only a single first image, the accuracy of the obtained abnormal driver behavior detection results can be further improved. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating Embodiment 1 of the driver abnormal behavior detection method of this application;

[0010] Figure 2 yes Figure 1 A detailed flowchart of the S11 process;

[0011] Figure 3 This is a schematic diagram of the second image;

[0012] Figure 4 This is a schematic diagram of the face region in the second image and the face region after edge expansion.

[0013] Figure 5 This is a flowchart illustrating Embodiment 2 of the driver abnormal behavior detection method of this application;

[0014] Figure 6 yes Figure 5 A detailed flowchart of the S21 process;

[0015] Figure 7 This is a schematic diagram of the detection network structure in this application;

[0016] Figure 8 This is a flowchart illustrating Embodiment 3 of the driver abnormal behavior detection method of this application;

[0017] Figure 9 This is a flowchart of Example 3 in this application;

[0018] Figure 10 This is a flowchart illustrating an embodiment of the training method for the detection network in this application;

[0019] Figure 11 This is a schematic diagram of the structure of an embodiment of the electronic device of this application;

[0020] Figure 12 This is a schematic diagram of the structure of an embodiment of the storage medium of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0022] The terms "first," "second," and "third" used in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments without conflict.

[0024] Figure 1 This is a flowchart illustrating Embodiment 1 of the driver abnormal behavior detection method of this application. It should be noted that if substantially the same result is obtained, this embodiment is not necessarily identical. Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment may include:

[0025] S11: Get multiple first images.

[0026] The first image contains at least one part of the driver.

[0027] The first image can be obtained by capturing the driver in a driving scenario using a camera, or it can be obtained by further processing the captured image (such as cropping, expanding edges, etc.). Taking the image captured by the camera as an example, multiple first images can be captured consecutively or at intervals. Multiple first images can be arranged into an image sequence, in which the first images are arranged in chronological order of capture time. The type of the first image can be RGB, IR, etc., depending on the type of camera.

[0028] The first image includes body parts of the driver that may exhibit or be associated with abnormal behavior. Possible abnormal behaviors could include yawning, smoking (with a cigarette in the mouth or hand), making or receiving phone calls, or taking the hands off the steering wheel. Therefore, body parts that may exhibit abnormal behavior could include the driver's face, hands, ears, etc.

[0029] See also Figure 2 If the first image contains the driver's face, and the first image is obtained by further processing the captured image, then acquiring multiple first images in S11 may include the following sub-steps:

[0030] S111: Multiple second images were obtained by taking pictures of the driver.

[0031] Multiple second images can be obtained by using a camera device to capture real-time images of the driver. For an example of a second image, please refer to [link / reference needed]. Figure 3 .

[0032] S112: Perform face detection on each second image to determine whether a face region exists in each second image.

[0033] If it exists, then execute S113.

[0034] S113: Expand the edges of the face region in the second image.

[0035] Edge expansion processing can extend the face region in the second image by a certain number of pixels to the surrounding areas (at least one direction: left, right, up, or down) to include the parts associated with abnormal behavior within the expanded face region, which is beneficial for subsequent detection.

[0036] For example, when a driver is answering a phone call, the area for answering the call consists of the hand, the phone, and the ear. Expanding the facial area can include this area. Similarly, when a driver is smoking, the area for smoking consists of the mouth and the cigarette. Expanding the facial area can include this area.

[0037] Figure 4 The number 1 in the image represents an example of a face region in the second image. Figure 4 The number 2 in the image represents an example of a face region in the second image that has undergone edge expansion processing.

[0038] S114: Extract the face region that has undergone edge expansion processing from the corresponding second image and use it as the first image.

[0039] S12: Use the detection network to detect abnormal behavior in each first image and obtain the detection result for each first image.

[0040] The detection network is a deep learning network with target (abnormal behavior) detection capabilities, i.e., a target detection network. It exhibits strong environmental adaptability and high robustness; therefore, compared to traditional image detection methods, this application utilizes a detection network to detect abnormal behavior in the first image, thereby improving the accuracy of the detection results.

[0041] The abnormal behavior detection of the first image in this application can be performed after the current first image is acquired (i.e., real-time detection, detecting one image at a time), or it can be performed after multiple first images are acquired (i.e., non-real-time detection, detecting multiple images at a time).

[0042] The detection results of the first image may include information on one or more abnormal behavior categories, and the information on each abnormal behavior category includes the probability of the existence of the abnormal behavior category and location information.

[0043] Location information can be the location information of areas where abnormal behavior may occur. This location information can be represented as coordinates or (coordinates, width, height), etc.

[0044] When there are multiple categories of abnormal behavior, the location information corresponding to the probabilities of different abnormal behavior categories can be the same or different. Specifically, a detection network can be used to obtain the location information of regions in the first image where abnormal behavior may exist, as well as the probability of each abnormal behavior category existing in those regions. The location information of these regions where abnormal behavior may exist can be directly used as the location information corresponding to the probability of each abnormal behavior category, that is, these regions where abnormal behavior may exist can be considered as the regions where each category of abnormal behavior may exist. Alternatively, the regions where abnormal behavior may exist can be further analyzed to determine the sub-regions where each category of abnormal behavior may exist within these regions.

[0045] For example, consider two types of abnormal behavior: smoking and making phone calls. The detection results for the first image include the location information of region A in the first image where abnormal behavior is suspected, as well as the probability of the smoking category and the probability of the phone call category. A can be directly considered as the region where smoking or phone calls might occur. Alternatively, A can be further analyzed to determine which specific sub-region within A is likely to contain smoking / phone calls.

[0046] It is understandable that, when the detection results of the first image include information on multiple abnormal behavior categories, this application can use the same detection network to detect multiple abnormal behaviors of the driver simultaneously. Compared with separate detection methods, this can improve detection speed and reduce the computational overhead required for detection.

[0047] S13: Based on the detection results of each first image, obtain the abnormal behavior detection results of the driver.

[0048] The detection results of multiple first images can be comprehensively analyzed, resulting in more accurate abnormal behavior detection results for drivers compared to cases where only one first image is used.

[0049] Furthermore, in other embodiments, after obtaining the abnormal behavior detection results of the driver, if the abnormal behavior detection results indicate that the driver is engaging in abnormal behavior, an alarm can be triggered to remind the driver to correct their abnormal behavior in a timely manner. Additionally, video within the time range corresponding to multiple first images can be recorded for subsequent analysis.

[0050] Through the implementation of this embodiment, this application utilizes a detection network to detect abnormal behavior from multiple first images, and comprehensively analyzes the abnormal behavior of the driver based on the detection results of multiple first images. Since the detection network is a deep learning network, it can improve the accuracy of the detection results of the first images compared to traditional image processing methods, thereby improving the accuracy of the abnormal behavior detection results of the driver obtained based on the detection results of the first images. Furthermore, compared to analyzing the abnormal behavior of the driver using only a single first image, it can further improve the accuracy of the obtained abnormal behavior detection results of the driver.

[0051] Figure 5 This is a flowchart illustrating Embodiment Two of the driver abnormal behavior detection method of this application. It should be noted that if substantially the same result is obtained, this embodiment is not necessarily identical. Figure 5 The illustrated process sequence is limited. This embodiment is a further extension of S12. Figure 1 As shown, this embodiment may include:

[0052] S21: For each first image, use a detection network to extract multiple anchor boxes from the first image.

[0053] The process involves using a detection network to extract features from the first image, obtaining a feature map corresponding to the first image, and then extracting multiple anchor boxes from the feature map. The regions corresponding to the anchor boxes in the first image can be areas in the first image where abnormal behavior may exist. The feature extraction involved in this application can include a series of operations such as convolution (Conv), pooling, and ReLU activation. The different feature extraction processes (first feature extraction, second feature extraction, and third feature extraction) described later may rely on different parameters, that is, the parameters of the corresponding convolution / pooling / activation operations may be different.

[0054] See also Figure 6 In one specific embodiment, extracting multiple anchor point boxes from the first image in step S21 may include the following steps:

[0055] S211: Perform first feature extraction on the first image to obtain the first feature map.

[0056] S212: The first feature map is processed using the receptive field mechanism to obtain the second feature map, and the second feature is extracted from the first feature map to obtain the third feature map.

[0057] S213: Fuse the second and third feature maps to obtain the fourth feature map.

[0058] During the fusion process, the weights of the first and second feature maps can be the same or different. If the weights are the same, the fusion can also be called an additive process, which involves adding corresponding positions in the second and third feature maps to obtain the fourth feature map.

[0059] S214: Based on the fourth feature map, obtain multiple anchor point boxes corresponding to the first image.

[0060] In one specific implementation, multiple anchor boxes can be extracted from the fourth feature map and used as multiple anchor boxes corresponding to the first image.

[0061] To better detect targets of different sizes in the first image, in another specific embodiment, at least one third feature extraction can be performed on the fourth feature map to obtain at least one fifth feature map; at least one anchor box is extracted from the fourth feature map and each fifth feature map respectively to obtain multiple anchor boxes corresponding to the first image. Each fifth feature map has a different resolution.

[0062] When multiple fifth feature maps are included, third feature extraction can be performed on the fourth feature map at different times to obtain fifth feature maps of different resolutions. Alternatively, a later fifth feature map can be obtained by extracting the third feature from a previous fifth feature map. In this case, multiple fifth feature maps can form a feature map pyramid.

[0063] S22: Perform abnormal behavior detection on each anchor box to obtain information on the abnormal behavior category of each anchor box.

[0064] The following combination Figure 7 The implementation process of S21-S22 will be illustrated with an example:

[0065] like Figure 7 As shown, the detection network may include a first feature extraction module, a receptive field module (RFB), a second feature extraction module, a fusion module, a third feature extraction module, a bounding box extraction module, and a detection module.

[0066] The RGB image (320*240*3) is fed into the detection network, and the first feature extraction module can be used to process the RGB image to obtain the first feature map (40*30*32).

[0067] The first feature map can be processed using the receptive field (RFB) module / mechanism to obtain the second feature map (40*30*32), and the second feature extraction module can be used to extract the second feature from the first feature map to obtain the third feature map (40*30*32).

[0068] The second and third feature maps can be added together using the fusion module to obtain the fourth feature map (40*30*32).

[0069] The third feature extraction module can be used to extract the third feature from the fourth feature map to obtain the fifth feature map a (20*15*64). Further extraction of the third feature from feature map a yields the fifth feature map b (10*8*128). Finally, extraction of the third feature from feature map b yields the fifth feature map c (5*4*128). The fifth feature maps a, b, and c constitute a feature map pyramid.

[0070] The bounding box extraction module can be used to extract anchor boxes of sizes 10*10, 16*16, and 24*24 on the fourth feature map, yielding a total of 40*30*3 = 3600 boxes. The same module can be used to extract anchor boxes of sizes 32*32 and 48*48 on the fifth feature map a, yielding a total of 20*15*2 = 600 boxes. The same module can be used to extract anchor boxes of sizes 64*64 and 96*96 on the fifth feature map b, yielding a total of 10*8*2 = 160 boxes. The same module can be used to extract anchor boxes of sizes 128*128, 192*192, and 256*256 on the fifth feature map c, yielding a total of 5*4*3 = 60 boxes. Therefore, a total of 3600 + 600 + 160 + 60 = 4420 anchor boxes are obtained.

[0071] The detection module can be used to detect abnormal behavior in 4420 anchor boxes and obtain information on the abnormal behavior category of each anchor box.

[0072] S23: Use the information of the abnormal behavior categories of some or all anchor boxes in the first image as the detection result of the first image.

[0073] Information on the abnormal behavior categories of all anchor boxes can be directly used as the detection result of the first image.

[0074] However, considering computational complexity and the potential presence of invalid anchor boxes (interference boxes), all anchor boxes can be filtered, and the abnormal behavior category information of the remaining anchor boxes after filtering can be used as the detection result of the first image. In this approach, S23 may include: filtering anchor boxes using non-maximum suppression based on the abnormal behavior category information of each anchor box; and using the abnormal behavior category information of the remaining anchor boxes after filtering as the detection result of the first image.

[0075] Non-maximum suppression (NMS) is used to filter anchor boxes, which means filtering out all interfering boxes in the anchor boxes based on the probability of abnormal behavior categories existing in the anchor boxes and the cross-union ratio between different boxes.

[0076] Figure 8 This is a flowchart illustrating Embodiment 3 of the driver abnormal behavior detection method of this application. It should be noted that if substantially the same result is obtained, this embodiment is not necessarily identical. Figure 8 The illustrated process sequence is limited. This embodiment is a further extension of S13. Anomaly detection results can include categories of existing abnormal behaviors. For example... Figure 8 As shown, this embodiment may include:

[0077] S31: For each abnormal behavior category, determine whether the probability of the first image having an abnormal behavior category meets the preset conditions.

[0078] A probability threshold can be preset for each type of abnormal behavior. A preset condition can be that the probability of an abnormal behavior type being present is greater than the corresponding probability threshold.

[0079] If the condition is met, then execute S32; if the condition is not met, then it can be determined that the abnormal behavior category does not exist in the first image.

[0080] S32: Determine the category of abnormal behavior in the first image.

[0081] S33: Determine whether the number of first images with abnormal behavior categories is greater than the preset number.

[0082] If it is greater than, then execute S34.

[0083] S34: Determine the category of abnormal behavior of the driver.

[0084] The first image of an abnormal behavior category that exceeds a preset number of images can be continuous or discontinuous.

[0085] The following three examples illustrate S31-S34 (the time period for one detection is the time period corresponding to 30 first images).

[0086] Example 1 (non-real-time, multiple images acquired before detection):

[0087] Obtain 30 first images to form an image sequence A

[30] =[m1,m2,…,m30];

[0088] Determine whether smoking exists in each of the numbers m1-m30;

[0089] If there are fewer than 10 instances of smoking in m1-m30, then it is determined that the driver did not smoke during the time period corresponding to m1-m30.

[0090] Update A

[30] = [m2, m3, ..., m31], determine whether there is smoking in m31. If there are more than 10 smoking photos in m2-m31, then determine that the driver was smoking in the time period corresponding to m2-m31; ..., and so on.

[0091] Example 2 (real-time, one image at a time):

[0092] Get the first image;

[0093] Determine if smoking is present in the first image;

[0094] If it exists, the number of the first images containing smoking is counted as 1; otherwise, it is counted as 0 (assuming it is 1).

[0095] Get the second image and determine if smoking is present in the second image;

[0096] If smoking exists, update the number of first images containing smoking to 2; otherwise, do not update; ... and so on, until the number of first images containing smoking is updated based on the judgment result of smoking in the 30th image.

[0097] Determine if the number of images showing smoking after the update is greater than 10. If it is, then determine that the driver was smoking during the time period corresponding to images 1-30.

[0098] Using the same method, it can be further determined whether the driver was smoking during the time period corresponding to images 2-31.

[0099] Example 3 (Applicable to both real-time and non-real-time, taking non-real-time as an example):

[0100] Combination Figure 9 To illustrate, obtain 30 first images to form an image sequence A

[30] =[m1,m2,…,m30];

[0101] Determine if smoking exists in m1;

[0102] If it does not exist, then it is directly determined that the driver did not smoke during the time period corresponding to A

[30] . If it does exist, then it is further determined whether smoking exists in m2-m30 respectively, and the number of people smoking in m1-m30 is determined according to the judgment result.

[0103] If there are more than 10 instances of smoking in m1-m30, then it is determined that the driver was smoking during the time period corresponding to A

[30] ; otherwise, it is determined that the driver was not smoking during the time period corresponding to A

[30] .

[0104] Furthermore, the detection network can be trained before being used in the above methods. Specifically, this can be done as follows:

[0105] Figure 10 This is a flowchart illustrating a first embodiment of the training method for the detection network in this application. It should be noted that if substantially the same result is obtained, this embodiment does not necessarily reflect that outcome. Figure 10 The process sequence shown is limited. Figure 10 As shown, this embodiment may include:

[0106] S41: Obtain a sample image containing the driver.

[0107] The sample images can be labeled with information on the actual abnormal behavior categories of the driver and the actual location information corresponding to the actual abnormal behavior categories.

[0108] The sample images containing drivers constitute a sample image set. This set can be composed of sample images collected under different environmental conditions (daytime, nighttime, sunny, cloudy, etc.) showing drivers engaging in various behaviors. For example, a sample image set could consist of 40,000 images of drivers (30 men and 30 women) in different environments, depicting them with cigarettes in their mouths, smoking, or making phone calls.

[0109] S42: Use a detection network to detect abnormal behavior in the sample images and obtain the detection results for each sample image.

[0110] The detection results of the sample images may include the predicted probability and predicted location information of at least one abnormal behavior category.

[0111] In one specific implementation, a detection network can be used to obtain multiple anchor bounding boxes corresponding to the sample image, and the detection network can be used to obtain the detection result of each anchor bounding box, that is, the predicted probability and predicted location information of at least one abnormal behavior category. For specific implementation details, please refer to the previous embodiments, which will not be repeated here.

[0112] The detection results of each anchor box (including the predicted probability and predicted location information of at least one abnormal behavior category) can be directly used as the detection results of the sample image. Alternatively, non-maximum suppression can be used to filter the anchor boxes, and the detection results of the remaining anchor boxes (including the predicted probability and predicted location information of at least one abnormal behavior category) can be used as the detection results of the sample image.

[0113] S43: Obtain the classification loss of the detection network and the localization loss of the detection network.

[0114] The classification loss of the detection network can be obtained based on the difference between the predicted probability of at least one abnormal behavior category for each anchor box in the detection results of the sample image and the corresponding true abnormal behavior category information. Similarly, the localization loss of the detection network can be obtained based on the difference between the predicted location information and the corresponding true location information of each anchor box in the detection results of the sample image.

[0115] The classification loss can be either the classification cross-entropy loss function or the L1 loss function, while the localization loss can be either the smooth L1 loss function or the L2 loss function. Compared to the L1 loss function, the multi-class cross-entropy loss function converges faster; compared to the L2 loss function, the smooth L1 loss function is less sensitive to outliers and anomalies, exhibits smaller gradient changes, and reduces the probability of overfitting.

[0116] The formula for calculating the classification cross-entropy loss function can be as follows:

[0117]

[0118] Among them, X i Y represents the predicted probability distribution of C abnormal behavior categories for the i-th anchor box. i This represents the true probability distribution (true abnormal behavior category information) of C types of abnormal behavior existing in the i-th anchor box, p ij y represents the probability that the i-th anchor box contains the j-th abnormal behavior category. ij This represents the true probability of the j-th abnormal behavior category (1 if the j-th abnormal behavior category exists in the i-th anchor box; 0 otherwise).

[0119] The formula for calculating the smooth L1 loss function can be as follows:

[0120]

[0121] Where x represents the difference between the predicted location information and the corresponding true location value.

[0122] S45: Adjust the parameters of the detection network based on the classification loss and localization loss.

[0123] The conditions for stopping training can be that the loss function converges, the number of training iterations reaches a threshold, or the training time reaches a threshold, etc.

[0124] Figure 11 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. Figure 11 As shown, the electronic device includes a processor 51 and a memory 52 coupled to the processor 51.

[0125] The memory 52 stores program instructions for implementing the methods of any of the above embodiments; the processor 51 executes the program instructions stored in the memory 52 to implement the steps of the above method embodiments. The processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 51 may be any conventional processor.

[0126] Figure 12 This is a schematic diagram of the structure of an embodiment of the storage medium of this application. Figure 12 As shown, the computer-readable storage medium 60 of this application embodiment stores program instructions 61, which, when executed, implement the methods provided in the above embodiments of this application. The program instructions 61 can form a program file and be stored in the computer-readable storage medium 60 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) or processor can execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned computer-readable storage medium 60 includes various media capable of storing program code, such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.

[0127] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for detecting abnormal driver behavior, characterized in that, include: Acquire multiple first images, wherein the first image contains at least one part of the driver; An abnormal behavior detection is performed on each of the first images using a detection network to obtain the detection result for each of the first images; Based on the detection results of each of the first images, the abnormal behavior detection results of the driver are obtained; The detection result of the first image includes information on one or more abnormal behavior categories; the step of using a detection network to perform abnormal behavior detection on each of the first images to obtain the detection result of each of the first images includes: for each first image, extracting multiple anchor boxes from the first image using the detection network; performing abnormal behavior detection on each anchor box to obtain information on the abnormal behavior category of each anchor box; using the information on the abnormal behavior categories of some or all of the anchor boxes as the detection result of the first image; and / or The first image contains the driver's face; acquiring multiple first images includes: taking pictures of the driver to obtain multiple second images; performing face detection on each second image to determine whether a face region exists in each second image; if it exists, expanding the edges of the face region in the second image; and extracting the expanded face region from the corresponding second image as the first image.

2. The method according to claim 1, wherein extracting multiple anchor point boxes from the first image comprises: The first feature is extracted from the first image to obtain a first feature map; The first feature map is processed using the receptive field mechanism to obtain the second feature map, and the second feature is extracted from the first feature map to obtain the third feature map; The second feature map and the third feature map are fused to obtain a fourth feature map; based on the fourth feature map, multiple anchor boxes corresponding to the first image are obtained.

3. The method according to claim 2, characterized in that, The process of fusing the second feature map and the third feature map to obtain the fourth feature map includes: Add the corresponding positions in the second feature map and the third feature map to obtain the fourth feature map; And / or, obtaining multiple anchor point boxes corresponding to the first image based on the fourth feature map includes: The fourth feature map is subjected to at least one third feature extraction to obtain at least one fifth feature map, wherein each fifth feature map has a different resolution; At least one anchor box is extracted from the fourth feature map and each of the fifth feature maps respectively to obtain multiple anchor boxes corresponding to the first image.

4. The method according to claim 1, characterized in that, Using information about the abnormal behavior of a portion of the anchor points as the detection result of the first image includes: Based on the information of the abnormal behavior category of each anchor box, the anchor boxes are filtered using a non-maximum suppression method; The information of the abnormal behavior category of the remaining anchor boxes after filtering is used as the detection result of the first image.

5. The method according to claim 1, characterized in that, Information for each of the abnormal behavior categories includes the probability of the existence of the abnormal behavior category and location information.

6. The method according to claim 5, characterized in that, The abnormal behavior detection results include the existing abnormal behavior categories. The abnormal behavior detection results for the driver, obtained based on the detection results for each of the first images, include: For each of the abnormal behavior categories, if the probability of the first image containing the abnormal behavior category meets a preset condition, then the first image is determined to contain the abnormal behavior category; if the number of first images containing the abnormal behavior category is greater than a preset number, then the driver is determined to contain the abnormal behavior category.

7. The method according to claim 1, characterized in that, Before performing abnormal behavior detection on each of the first images using the detection network to obtain the detection result for each of the first images, the following steps are included in training the detection network: Obtain a sample image containing the driver; The detection network is used to detect abnormal behavior in the sample images to obtain the detection result for each sample image; The classification loss and localization loss of the detection network are obtained. The parameters of the detection network are adjusted based on the classification loss and the localization loss.

8. An electronic device, characterized in that, Includes a processor and a memory connected to the processor, wherein, The memory stores program instructions; The processor is used to execute the program instructions stored in the memory to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Smoking detection method, storage medium and computer

    CN108629282A

  • Abnormal behavior detection method and device and vehicle-mounted equipment

    CN109886209A

  • Fatigue driving detection method and device, computer equipment and storage medium

    CN111645695A