Smoking behavior recognition method, smoking detection model, device, vehicle and medium

By using infrared and visible light image fusion technology, the fusion weights and actual distances of human images are determined to generate a smoking behavior recognition model. This solves the problem of low recognition rate of smoking behavior in complex environments and achieves higher recognition accuracy and lower false alarm rate.

CN116486383BActive Publication Date: 2025-11-25GREAT WALL MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310145939.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-11-25
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

In complex environments, the recognition rate of smoking behavior inside a car cabin is low, and false alarms are prone to occur.

Method used

By employing infrared and visible light image fusion technology, and determining the fusion weights of the person's image and the actual distance between the person's mouth and the cigarette, fused image features are generated, and a pre-trained classification model is used to identify smoking behavior.

Benefits of technology

It improves the accuracy of smoking behavior recognition in complex environments and reduces the false alarm rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486383B_ABST
    Figure CN116486383B_ABST
Patent Text Reader

Abstract

The present disclosure provides a smoking behavior recognition method, a smoking detection model, a device, a vehicle and a medium, and relates to the technical field of image recognition. The method comprises the following steps: determining a fusion weight of a person image in a position of a driver in a vehicle cabin according to an image display parameter, determining an actual distance between a mouth of the person in the person image and a cigarette, generating a fusion image feature according to the fusion weight and the actual distance, classifying a smoking behavior of the driver according to the fusion image feature, and obtaining a recognition result of whether the smoking behavior exists. The present disclosure can improve the detection accuracy of whether the smoking behavior of the driver exists, and is beneficial to reduce the false detection rate of the smoking behavior.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image recognition, and particularly relates to a smoking behavior recognition method, a smoking detection model, an apparatus, a vehicle and a computer readable storage medium. BACKGROUND

[0002] Artificial intelligence (AI) visual recognition technology has been applied in various industries. By linking AI capability layers, a camera can have the ability to automatically identify safety hazards and can link a sound and light alarm to issue a reminder to achieve closed-loop management of hazards.

[0003] With the rapid development of intelligent automobiles, the automobile cabin is the most important part of human-machine interaction in automobiles. An AI visual recognition technology-based method for monitoring dangerous behavior in an automobile cabin using a single camera video stream is proposed. Since the video stream used for monitoring dangerous behavior in the automobile cabin is collected by a single camera, the recognition rate of dangerous behavior in complex environment scenarios (for example, bright, dark, shadow, mottled light, and other scenarios in which other reflective objects in the automobile cabin reflect light or light spots onto the faces of drivers and passengers, resulting in light spots on the faces) is low, and false positives of dangerous behavior are prone to occur.

[0004] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0005] The present disclosure provides a smoking behavior recognition method, a smoking detection model, an apparatus, a vehicle and a computer readable storage medium, which can improve the recognition rate of dangerous behavior in an automobile cabin in a complex environment scenario and reduce the false positive rate of dangerous behavior.

[0006] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.

[0007] According to one aspect of the present disclosure, a smoking behavior recognition method is provided, the smoking behavior recognition method comprising: determining a fusion weight of a person image according to an image display parameter; wherein the person image is an image of a position of a driver in a cabin; determining an actual distance between a mouth of a person in the person image and a cigarette; generating a fused image feature according to the fusion weight and the actual distance; and classifying a smoking behavior of the driver according to the fused image feature to obtain a recognition result of whether there is smoking.

[0008] Optionally, the image of the person includes an infrared image and a visible light image; before the step of determining the fusion weight of the image of the person according to the image display parameter, the smoking behavior recognition method further includes: acquiring a first video stream of a position where the driver is located in a vehicle cabin, and performing frame extraction processing on the first video stream to obtain the infrared image, the first video stream being collected by an infrared camera; acquiring a second video stream of an overall environment inside the vehicle cabin, and performing frame extraction processing on the second video stream to obtain a vehicle cabin environment image, the second video stream being collected by a visible light camera; and cutting out an image of the position from the vehicle cabin environment image to obtain the visible light image.

[0009] Optionally, the image display parameter includes an image gray value; and the step of determining the fusion weight of the image of the person according to the image display parameter includes: determining a first image gray average value and a first image gray standard deviation of the infrared image according to an image gray value of the infrared image; determining a second image gray average value and a second image gray standard deviation of the visible light image according to an image gray value of the visible light image; and performing normalization processing on the first image gray average value, the first image gray standard deviation, the second image gray average value and the second image gray standard deviation to obtain the fusion weight.

[0010] Optionally, the step of determining the actual distance between the mouth of the person and the cigarette in the image of the person includes: identifying a center point position of the mouth of the person and a center point position of the cigarette in the image of the person by using a pre-trained target detection model; calculating a distance between the center point position of the mouth of the person and the center point position of the cigarette to obtain a center point distance; and determining the center point distance as the actual distance.

[0011] Optionally, the step of classifying the smoking behavior of the driver according to the fusion image feature to obtain the recognition result of whether there is smoking includes: inputting the fusion image feature into a pre-trained classification model to classify the smoking behavior of the driver and obtain the recognition result of whether there is smoking.

[0012] Optionally, before the step of classifying the smoking behavior of the driver according to the fusion image feature to obtain the recognition result of whether there is smoking, the smoking behavior recognition method further includes: acquiring an infrared face image and a visible light face image; determining a sample image gray average value and a sample image gray standard deviation of each of the infrared face image and the visible light face image according to an image gray value to obtain a sample fusion weight; determining a distance value pair between the mouth of a person and a cigarette in each of the infrared face image and the visible light face image; generating a sample fusion image feature according to the sample fusion weight and the distance value pair; and training the classification model by using the sample fusion image feature.

[0013] According to another aspect of the present disclosure, there is provided a smoking detection model, comprising:

[0014] a fusion weight module configured to determine a fusion weight of both the infrared image and the visible light image according to an image display parameter in a case where a human face exists in the infrared image and / or the visible light image;

[0015] a pre-trained target detection model configured to determine an actual distance between a human mouth and a cigarette in each of the infrared image and the visible light image;

[0016] a feature fusion module configured to generate a fused image feature according to the fusion weight and the actual distance;

[0017] a pre-trained classification model configured to classify a smoking behavior of a human in the infrared image and the visible light image according to the fused image feature to obtain a recognition result of whether smoking exists.

[0018] According to still another aspect of the present disclosure, there is provided a smoking behavior recognition device, comprising:

[0019] a weight calculation module configured to determine a fusion weight of a human image according to an image display parameter;

[0020] a distance calculation module configured to determine an actual distance between a human mouth and a cigarette in the human image;

[0021] a feature fusion module configured to generate a fused image feature according to the fusion weight and the actual distance;

[0022] a behavior recognition module configured to classify a smoking behavior of the driver according to the fused image feature to obtain a recognition result of whether smoking exists.

[0023] According to still another aspect of the present disclosure, there is provided a vehicle, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the smoking behavior recognition method as described in the above embodiments, or implement the smoking behavior recognition method.

[0024] According to still another aspect of the present disclosure, there is provided a computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the smoking behavior recognition method as described in the above embodiments, or implement the smoking behavior recognition method.

[0025] The smoking behavior recognition method, the smoking detection model, the device, the vehicle and the computer readable storage medium provided by the embodiment of the present disclosure have the following technical effects.

[0026] The present disclosure determines the actual distance between the mouth of the person in the person image and the cigarette in the vehicle cabin according to the image display parameters, generates a fusion image feature according to the fusion weight and the actual distance, classifies the smoking behavior of the driver according to the fusion image feature, and obtains the recognition result of whether the driver smokes. The technical scheme increases the basis data for identifying whether the driver smokes by taking the fusion weight of the person image related to the image display parameters and the actual distance between the mouth of the person in the person image and the cigarette as the fusion image feature, and can improve the detection accuracy of whether the driver smokes, thereby reducing the false detection rate of the smoking behavior.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0028] The accompanying drawings incorporated in the specification and forming a part of the specification illustrate the embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0029] Figure 1 A flowchart of model training in an exemplary embodiment of the present disclosure is shown;

[0030] Figure 2 A flowchart of the smoking behavior recognition method in an exemplary embodiment of the present disclosure is shown;

[0031] Figure 3 A schematic diagram of the vehicle cabin environment image is shown;

[0032] Figure 4 A schematic diagram of the first labeled image is shown;

[0033] Figure 5 A flowchart of the smoking behavior detection using the smoking detection model is shown;

[0034] Figure 6 A structure diagram of the smoking detection model of the present disclosure is shown;

[0035] Figure 7 A processing flowchart of the smoking detection model of the present disclosure is shown;

[0036] Figure 8 Fig. 1 shows a structural schematic diagram of a smoking behavior recognition device provided by an embodiment of the present disclosure;

[0037] Figure 9 Fig. 2 shows a structural schematic diagram of a vehicle provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] For the purposes of the present disclosure, technical solutions and advantages, the following will be further described in detail with reference to the drawings.

[0039] The following description refers to the accompanying drawings. Unless otherwise indicated, like numbers in the various drawings of the accompanying drawings denote like or similar elements. The following detailed description does not provide design alternatives that are not expressly described in the following example embodiments. Rather, the following description provides example embodiments in accordance with some aspects of the present disclosure as detailed in the appended claims.

[0040] It should be noted that the above-described drawings are only schematic representations of the processes included in the method according to the example embodiments of the present disclosure, and are not intended to limit the purposes. It is easy to understand that the processes shown in the above-described drawings do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0041] With the rapid development of intelligentization of automobiles, the automobile cabin as the most important part of human-computer interaction in the automobile, in order to ensure the safety in the automobile cabin, it is necessary to monitor the dangerous behavior (such as smoking behavior) in the automobile cabin. In the related art, based on the AI vision recognition technology, a method for monitoring the dangerous behavior in the automobile cabin based on the video stream of a single camera is proposed. Since the video stream used for monitoring the dangerous behavior in the automobile cabin is collected by a single camera, the recognition rate of the dangerous behavior in a complex environment scene (such as the case where the light or light spot is reflected on the face of the driver or passenger by other reflective objects in the automobile cabin under bright, dark, shadow, mottled light, etc. Scene, resulting in light spots on the face) is low, and false positives of dangerous behavior are prone to occur.

[0042] Based on the problems existing in the above-mentioned related art, the present disclosure provides a smoking behavior recognition method, a smoking detection model, a device, a vehicle and a computer readable storage medium, to solve the problem of low recognition rate of smoking behavior in a complex environment scene and prone to false positives of dangerous behavior.

[0043] The monitoring of smoking behavior in the automobile cabin in a complex environment scene can be identified by a smoking detection model. As shown in Figure 1 Figure 1 ​A flow chart of model training in an example embodiment of the present disclosure is shown. For the identification of smoking behavior, at least two models need to be trained, namely a face recognition model and a smoking detection model, and the training process is as follows:

[0044] Data collection, i.e. collecting OMS (Occupant Monitoring System) data and DMS (Driver Monitoring System) data. Among them, the OMS data refers to the video data of the overall environment inside the car cabin. The OMS data is collected by an RGB (Red Green Blue) camera. The monitoring object of the OMS system is the passenger. The RGB camera is arranged at multiple positions in the cabin. The RGB camera can capture all positions inside the car cabin, that is, the RGB camera can not only capture the passenger but also capture the driver. The DMS data refers to the video data of the position of the driver in the car cabin. The DMS data is collected by an IR (Infrared) camera. The monitoring object of the DMS system is the driver. The IR camera is arranged at the left A-pillar position of the car. The IR camera mainly collects video information of the area where the driver's seat is located, that is, in the case that a person sits on the driver's seat, the IR camera will capture the person on the driver's seat. Among them, the IR camera is used for night vision monitoring. The focal position of the lens of an ordinary camera will change under the condition of infrared light at night, so that the image will become blurred and needs to be adjusted to be clear. The focal point of the lens of the IR camera is consistent under both infrared light and visible light.

[0045] Collecting OMS data and DMS data can be performed simultaneously. An example collection process is as follows: replace the driver with different (age, skin color, gender, fat, height, etc.) sitting in the driver's seat to simulate a normal driving scene, and let the driver adjust the driver's seat to the appropriate driving state, and then collect OMS data through the RBG camera and collect DMS data through the IR camera. After collecting OMS data and DMS data, data processing is performed.

[0046] For data processing, a plurality of visible light images are obtained from the OMS data, and a plurality of infrared images are obtained from the DMS data. The obtained visible light images and infrared images can be understood as sample images, and the visible light images and infrared images need to contain faces. A part of the visible light images and infrared images also contain cigarettes, and another part of the visible light images and infrared images do not contain cigarettes.

[0047] After obtaining the sample images, they are labeled by marking face bounding boxes and cigarette bounding boxes, resulting in labeled sample images. These labeled sample images are then normalized to obtain normalized sample images. Next, the labeled sample images are used to train a face recognition model, and the normalized sample images are used to train a smoking detection model. Once the face recognition and smoking detection models are trained, they can be used in conjunction to detect whether a person is smoking.

[0048] During model training, OMS and DMS data are collected through simulated driving scenarios. The purpose of collecting OMS and DMS data is to obtain visible light and infrared images of people with faces and cigarettes. Additionally, visible light and infrared images of people with faces and cigarettes can be acquired in other scenarios. For example, an IR camera can be placed at one location indoors to collect video data of people smoking and non-smoking, from which infrared images of smoking and non-smoking individuals can be obtained. An RGB camera can be placed at another location indoors to collect panoramic video data, which can also be used to obtain visible light images of people smoking and non-smoking individuals. In this way, visible light and infrared images of people with faces and cigarettes can be acquired in various scenarios.

[0049] The following are embodiments of the smoking behavior recognition method provided in this disclosure.

[0050] in, Figure 2 A flowchart illustrating a smoking behavior recognition method in an exemplary embodiment of this disclosure is shown. Figure 2 As shown, an embodiment of the smoking behavior recognition method provided by this disclosure is applied in a driving scenario. The executing entity can be a vehicle controller or an in-vehicle terminal, and includes the following schemes:

[0051] S110: Determine the fusion weight of the person image based on the image display parameters.

[0052] In an exemplary embodiment, the person image is an image of the driver's location within the vehicle cabin. While the driver is driving the vehicle, an image of the driver's location within the vehicle cabin is acquired to obtain the person image. The person image is then input into a pre-trained face recognition model, which identifies whether a face exists in the person image. If a face is detected, the fusion weights of the person image are calculated based on the image display parameters.

[0053] The FaceNet algorithm can be selected as the face recognition model, a main network architecture of which is Inception-ResNetV1, and the face recognition model can obtain the prediction of the face frame, the mouth frame and the cigarette frame in the image, that is, the prediction of the face frame, the mouth frame and the cigarette frame in the image. The face recognition model is trained by using infrared sample images and visible light sample images. The infrared sample images and the visible light sample images contain faces. Among them, part of the infrared sample images contain not only the mouth of the person but also the cigarette, and the other part of the infrared sample images contain the mouth of the person but do not contain the cigarette. Similarly, part of the visible light sample images contain not only the mouth of the person but also the cigarette, and the other part of the visible light sample images contain the mouth of the person but do not contain the cigarette.

[0054] Before training the face recognition model, the faces in the infrared sample images and the visible light sample images need to be labeled, and then the labeled infrared sample images and the labeled visible light sample images are input into the Inception-ResNetV1 architecture. The architecture outputs the face feature vector, and then the face recognition model is trained by using the face feature vector until the model converges, and the face recognition model is trained.

[0055] The image display parameter includes an image gray value, which is the gray value of a pixel point in the image, that is, the image display parameter can be represented by the image gray value. The image gray average value and the image gray standard deviation of the person image are calculated, and after the image gray average value and the image gray standard deviation are normalized, a first value corresponding to the image gray average value and a second value corresponding to the image gray standard deviation are obtained. The first value and the second value are both between 0 and 1, and the first value and the second value represent the fusion weight of the person image, and the fusion weight is related to the image gray value. Among them, the first value is represented as p, the second value is represented as q, and the fusion weight is represented as (p, q), and p+q=1.

[0056] In a possible implementation, the person image includes an infrared image and a visible light image, and before the step of determining the fusion weight of the person image according to the image display parameter, the smoking behavior recognition method further includes the following scheme:

[0057] A first video stream of a position where the driver is located in the vehicle cabin is acquired, and frame extraction processing is performed on the first video stream to obtain the infrared image, and the first video stream is collected by an infrared camera;

[0058] A second video stream of an overall environment inside the vehicle cabin is acquired, and frame extraction processing is performed on the second video stream to obtain a vehicle cabin environment image, and the second video stream is collected by a visible light camera;

[0059] An image at the position is intercepted from the vehicle cabin environment image to obtain the visible light image.

[0060] The person image is acquired in a driving scene. A first video stream of a position of a driver in a vehicle cabin is acquired. The first video stream corresponds to DMS data, that is, the first video stream is collected by an IR camera. Frame extraction is performed on the first video stream to obtain a plurality of first to-be-processed images. Each first to-be-processed image is cropped to obtain a plurality of first cropped images. Each first cropped image is subjected to image grayscale processing to obtain a plurality of infrared images. Each infrared image has the same size, for example, the size is 800*1280.

[0061] A second video stream of an overall environment inside the vehicle cabin is acquired. The second video stream corresponds to OMS data, that is, the second video stream is collected by an RGB camera. Frame extraction is performed on the second video stream to obtain a plurality of second to-be-processed images. Each second to-be-processed image is cropped to obtain a plurality of vehicle cabin environment images. As shown in Figure 3 Figure 3 A schematic diagram of a vehicle cabin environment image is shown. Figure 3 P1 in the figure represents a vehicle cabin environment image, P2 represents an image of a position of a driver in the vehicle cabin, and a region of P1 other than P2 is an image of other things in the vehicle cabin, for example, the image of the other things includes rear seats, a co-driver seat, and the like.

[0062] For each vehicle cabin environment image, an image of the position of the driver is intercepted from each vehicle cabin environment image to obtain a plurality of second pre-processed images, for example, P2 in the figure. Figure 3 Then, each second pre-processed image is subjected to image grayscale processing to obtain a plurality of visible light images. Each visible light image has the same size, for example, the size is 1920*1080.

[0063] Because the positions of the RGB camera and the IR camera are different, the positions corresponding to the collected video data are also different, so that the person image of the same driver photographed at different positions can be acquired.

[0064] After the infrared image and the visible light image are acquired, the infrared image and the visible light image are sequentially input to a face recognition model. If the face recognition model recognizes that a face exists in the infrared image and / or the visible light image, a fusion weight of the infrared image and the visible light image is calculated according to an image display parameter.

[0065] In a possible implementation manner, the calculating the fusion weight of the person image according to the image display parameter includes the following scheme:

[0066] ​According to the image gray value of the infrared image, a first image gray mean value and a first image gray standard deviation of the infrared image are determined;

[0067] According to the image gray value of the visible light image, a second image gray mean value and a second image gray standard deviation of the visible light image are determined;

[0068] The first image gray mean value, the first image gray standard deviation, the second image gray mean value and the second image gray standard deviation are normalized to obtain the fusion weight.

[0069] The image gray mean value is calculated by formula (1), and the image gray standard deviation is calculated by formula (2), and formula (1) and formula (2) are as follows:

[0070] (1)

[0071] (2)

[0072] In the formula, x(i,j) indicates the gray value of the pixel point in the image, arg indicates the image gray mean value, (i,j) indicates the image gray standard deviation, A indicates the width of the image, and B indicates the length of the image. σ

[0073] Based on formula (1) and formula (2), the first image gray mean value and the first image gray standard deviation of the infrared image are calculated according to the image gray value of the infrared image, the first image gray mean value is represented as ir arg , and the first image gray standard deviation is represented as ir σ . Wherein, ir arg and ir σ represent the image weight value of the infrared image.

[0074] Similarly, based on formula (1) and formula (2), the second image gray mean value and the second image gray standard deviation of the visible light image are calculated according to the image gray value of the visible light image, the second image gray mean value is represented as rgb arg , and the second image gray standard deviation is represented as rgb σ . Wherein, rgb arg and rgb σ represent the image weight value of the visible light image.

[0075] ​After obtaining the image weight values of the infrared image and the visible light image respectively, the obtained image weight values are linearly fused to obtain the fusion weight corresponding to the infrared image and the visible light image, that is, the fusion weight includes the image weight values of the infrared image and the visible light image respectively. Since the infrared image and the visible light image are images of the same person collected at different positions, for the same person, the infrared image and the visible light image correspond to images at two angles or two states, and the fusion weight corresponding to the infrared image and the visible light image obtained is dynamic, that is, the dynamic fusion weight. The image gray average value and the image gray standard deviation corresponding to the infrared image and the visible light image are (ir arg , ir σ , rgb arg , rgb σ ), after normalization processing on (ir arg , ir σ , rgb arg , rgb σ ), the fusion weight is (a, b, c, d), wherein ir arg corresponds to a, ir σ corresponds to b, rgb arg corresponds to c, and rgb σ corresponds to d, a+b+c+d=1.

[0076] Since the image weight values of the infrared image and the visible light image are calculated by the image gray values, and the fusion weight is obtained by fusing the image weight values of the infrared image and the visible light image, it is equivalent to realizing the fusion of the infrared image and the visible light image, so as to realize the complementary advantages of the infrared image and the visible light image, and to avoid the interference of light on the face image in a complex environment scene.

[0077] S120: determining the actual distance between the mouth of the person in the person image and the cigarette.

[0078] After obtaining the person image, the actual distance between the mouth of the person in the person image and the cigarette is calculated, which is a distance value, denoted as k, which is part of the feature of the following fusion image.

[0079] In one possible implementation, the determination of the actual distance between the mouth of the person in the person image and the cigarette includes the following scheme:

[0080] A pre-trained target detection model is used to identify the center point positions of the mouth of the person and the center point positions of the cigarette in the person image;

[0081] The distance between the center point positions of the mouth of the person and the center point positions of the cigarette is calculated to obtain a center point distance;

[0082] The center point distance is determined as the actual distance.

[0083] YOLOV5 algorithm can be selected as the target detection model, which is pre-trained and used to identify the center point positions of cigarettes and mouths in images. The training of the above target detection model includes the following schemes:

[0084] A first labeled image carrying cigarette bounding boxes and mouth bounding boxes, and a second labeled image carrying mouth bounding boxes are obtained. The target detection model is trained by sampling the first labeled image and the second labeled image. The first labeled image is an image containing cigarettes obtained after classification of a person sample image, and the second labeled image is an image not containing cigarettes obtained after classification of a person sample image. The person sample image contains a face, specifically a mouth of a person.

[0085] First, the person sample images are classified according to smoking and non-smoking to obtain smoking images and non-smoking images. The smoking images contain cigarettes, and the non-smoking images do not contain cigarettes. The cigarettes and mouths in the smoking images are frame labeled to obtain cigarette bounding boxes and mouth bounding boxes, thereby generating the first labeled image carrying the cigarette bounding boxes and the mouth bounding boxes. That is, the first labeled image is an image containing cigarettes obtained after classification of a person sample image, which can also be understood as a smoking labeled image. Figure 4 As shown in Figure 4 a schematic diagram of the first labeled image is shown. Wherein, 100 represents the mouth, 101 represents the mouth bounding box, 200 represents the cigarette, and 201 represents the cigarette bounding box. Similarly, the mouth in the non-smoking image is frame labeled to obtain the mouth bounding box, thereby generating the second labeled image carrying the mouth bounding box. That is, the second labeled image is an image not containing cigarettes obtained after classification of a person sample image, which can also be understood as a non-smoking labeled image.

[0086] The first labeled image and the second labeled image are obtained, and the target detection model is iteratively trained using the first labeled image and the second labeled image until the model converges, thereby completing the training of the target detection model.

[0087] The target detection model obtains the center point positions of the mouth and the cigarette in the infrared image. The center point positions of the mouth and the cigarette in the infrared image can be represented by the coordinates of the pixel points. The distance between the coordinates of the pixel points corresponding to the center point positions of the mouth and the cigarette in the infrared image is calculated to obtain a center distance, which is taken as the actual distance between the mouth and the cigarette in the infrared image. Similarly, the target detection model obtains the center point positions of the mouth and the cigarette in the visible light image. The center point positions of the mouth and the cigarette in the visible light image can be represented by the coordinates of the pixel points. The distance between the coordinates of the pixel points corresponding to the center point positions of the mouth and the cigarette in the visible light image is calculated to obtain a center distance, which is taken as the actual distance between the mouth and the cigarette in the visible light image.

[0088] For the case that the image of the person includes the infrared image and the visible light image, the infrared image and the visible light image are input to the target detection model. The infrared image and the visible light image are images in which the presence of a face is identified by the face recognition model. After the face recognition model identifies the presence of a face in the infrared image and the visible light image, it also predicts a predicted face frame, a predicted mouth frame in each of the infrared image and the visible light image. If a cigarette is present in the image, a predicted cigarette frame of the cigarette can also be predicted.

[0089] Since there are many smoking postures such as pinching a cigarette, holding a cigarette, and holding a cigarette with a hand, and some cigarette tips have a cylindrical, rectangular, or other shape, the actual distance between the mouth and the cigarette in the image is set to determine whether the person has a smoking behavior. The calculation formula of the actual distance between the mouth and the cigarette is as follows:

[0090] (3)

[0091] In formula (3), d represents the distance, (x, y) represents the coordinates of the center point position of the predicted mouth frame in the image, x m , y m (x, y) represents the coordinates of the center point position of the predicted cigarette frame in the image, x s , y s (x, y) represents the coordinates of the center point position of the predicted face frame in the image, l represents the width of the predicted face frame in the image.

[0092] The infrared image and the visible light image are input into a target detection model, the target detection model obtains coordinates of a pixel point corresponding to a center point position of a predicted mouth frame, coordinates of a pixel point corresponding to a center point position of a predicted cigarette frame, and a width of a predicted face frame in each of the infrared image and the visible light image. Then, an actual distance between a mouth of a person and a cigarette in the infrared image is calculated by using formula (3) and is denoted as e, and an actual distance between the mouth of the person and the cigarette in the visible light image is calculated by using formula (3) and is denoted as f.

[0093] S130: generating a fusion image feature according to the fusion weight and the actual distance.

[0094] For the case that the person image is not limited, the calculated fusion weight is (p, q), and the actual distance is k. Then, the fusion weight and the actual distance are fused to obtain a fusion image feature, denoted as W1, W1=(p, q, k).

[0095] For the case that the person image is limited, that is, the person image includes the infrared image and the visible light image, the calculated fusion weight is (a, b, c, d), the actual distance between the mouth of the person and the cigarette in the infrared image is e, and the actual distance between the mouth of the person and the cigarette in the visible light image is f. Then, the fusion weight and the two actual distances are fused to obtain a fusion image feature, denoted as W2, W2=(a, b, c, d, e, f).

[0096] S140: classifying the smoking behavior of the driver according to the fusion image feature to obtain an identification result of whether there is smoking.

[0097] After the fusion image feature is obtained, the smoking behavior of the driver is classified according to the fusion image feature to obtain an identification result of whether there is smoking. If it is obtained through the identification result that the driver has the smoking behavior, a prompt information is sent to alarm the smoking behavior of the driver.

[0098] In a possible implementation, the classifying the smoking behavior of the driver according to the fusion image feature to obtain an identification result of whether there is smoking includes the following solutions:

[0099] The fusion image feature is input into a pre-trained classification model to classify the smoking behavior of the driver to obtain the identification result of whether there is smoking.

[0100] After the fusion image feature is obtained, the fusion image feature is input into a classification model, and the classification model outputs a recognition result, the recognition result including a probability value of the driver smoking, and whether the driver has a smoking behavior is represented by the recognition result. For example, when the probability value is greater than or equal to a preset value, it is represented that the driver has a smoking behavior, and when the probability value is less than the preset value, it is represented that the driver has a smoking behavior.

[0101] In a possible implementation, before the step of classifying the smoking behavior of the driver according to the fusion image feature to obtain a recognition result of whether there is smoking, the smoking behavior recognition method further includes the following scheme, that is, training of a classification model:

[0102] An infrared face image and a visible light face image are obtained.

[0103] According to image gray values, sample image gray average values and sample image gray standard deviations of the infrared face image and the visible light face image are determined, and sample fusion weights are obtained.

[0104] Distances between mouths of persons and cigarettes in the infrared face image and the visible light face image are determined, and a distance value pair is obtained.

[0105] According to the sample fusion weights and the distance value pair, a sample fusion image feature is generated.

[0106] The sample fusion image feature is used to train the classification model.

[0107] An infrared face image captured by an IR camera and a visible light face image captured by an RGB camera are obtained, and the infrared face image and the visible light face image contain a face. Part of the infrared face image contains a cigarette in addition to a mouth of a person, and another part of the infrared face image contains only the mouth of the person. Similarly, part of the visible light face image contains a cigarette in addition to the mouth of the person, and another part of the visible light face image contains only the mouth of the person.

[0108] According to the above formula (1) and formula (2), sample image gray average values and sample image gray standard deviations of the infrared face image and the visible light face image are calculated, and then the sample image gray average values and the sample image gray standard deviations of the infrared face image and the visible light face image are normalized to obtain sample fusion weights. For example, the sample image gray average value and the sample image gray standard deviation of the infrared face image are represented as ir1 arg and ir1 σ , respectively, and the sample image gray average value and the sample image gray standard deviation of the visible light face image are represented as rgb1 arg and rgb1 σ, for ir1 arg ir1 σ rgb1 arg and rgb1 σ After normalization, the sample fusion weights are obtained, which are (a1,b1,c1,d1) and a1+b1+c1+d1=1.

[0109] The distance between the mouth of the person and the cigarette in the infrared face image and the visible light face image is calculated according to the formula (3) above. The distance between the mouth of the person and the cigarette in the infrared face image is denoted as e1, and the distance between the mouth of the person and the cigarette in the visible light face image is denoted as f1. The distance value pair is (e1, f1).

[0110] After obtaining the sample fusion weight W3 and the distance value pair, the two are fused to obtain the sample fusion image features, denoted as W3, W3=(a1,b1,c1,d1,e1,f1).

[0111] The classification model is trained using sample fusion weights W3 until it converges, indicating that the classification model training is complete. The classification model employs the random forest algorithm.

[0112] In complex environments, light spots can appear on faces. For example, the cross-section of a cigarette butt is circular or rectangular. When light shines on a face, it can create light spots resembling the shape of a cigarette butt, interfering with the identification of whether someone is smoking and leading to misidentification. When training a classification model, since infrared and visible light face images are acquired from different locations, and infrared images contain infrared light while visible light images contain visible light, using the fused image features generated from these two images as input to the classification model is equivalent to using both infrared and visible light face images as input. This leverages the complementary nature of the infrared and visible light face images, avoiding misidentification of smoking behavior caused by cigarette light spots projected onto faces in complex environments.

[0113] The embodiment provided in the present disclosure adopts the technical scheme of determining the fusion weight of the figure image of the position of the driver in the cabin according to the image display parameter, determining the actual distance between the mouth of the figure in the figure image and the cigarette, generating the fusion image feature according to the fusion weight and the actual distance, and classifying the smoking behavior of the driver according to the fusion image feature to obtain the recognition result of whether the smoking exists. The fusion weight of the figure image related to the image display parameter and the actual distance between the mouth of the figure in the figure image and the cigarette are used as the fusion image feature, the basis data for identifying whether the driver has the smoking behavior is increased, the detection accuracy of whether the driver has the smoking behavior can be improved, and the false detection rate of the smoking behavior is reduced.

[0114] When the figure image includes the infrared image and the visible light image, the fusion image feature not only includes the fusion weight of the infrared image and the visible light image related to the image gray value, but also includes the actual distance between the mouth of the figure in the infrared image and the cigarette and the actual distance between the mouth of the figure in the visible light image and the cigarette. The basis data for identifying whether the driver has the smoking behavior is further increased, which is equivalent to fusing the infrared image and the visible light image, realizing the direct correlation between the complex environment scene and the smoking behavior, and obtaining the correlation between the smoking behavior and the complex environment scene. The fusion image feature related to the smoking behavior is used for smoking behavior recognition, the situation of predicting the smoking behavior from the infrared image and the visible light image is observed from the front, the detection accuracy of whether the driver has the smoking behavior is improved, and the false detection of the smoking behavior is reduced. When the classification model is applied to the smoking behavior recognition in the cabin of the car, the recognition rate of the smoking behavior in the cabin of the car in the complex environment scene is improved, and the false positive rate of the smoking behavior is reduced.

[0115] The following is an embodiment of the smoking detection model provided in the present disclosure.

[0116] After the smoking detection model and the face recognition model are trained, the smoking detection model and the face recognition model are used in cooperation to detect the smoking behavior. As shown in Figure 5 Figure 5 A flowchart for detecting the smoking behavior using the smoking detection model is shown, and the detection process is as follows:

[0117] ​The IR camera and the RGB camera arranged at different positions are used to collect the area infrared image and the area visible light image of the same person, and then the collected area infrared image and area visible light image are preprocessed, that is, the infrared image of the person is cropped from the area infrared image and the infrared image of the person is cropped from the area visible light image, and the image grayscale processing is performed. Then, whether there is a face in the two images after image grayscale processing is identified by a face recognition model, if there is a face in at least one of the two images, the two images after image grayscale processing are input into a smoking detection model, and the smoking detection model detects whether the person has a smoking behavior and outputs a detection result.

[0118] As shown in Figure 6 , Figure 6 A structural diagram of the smoking detection model of the present disclosure is shown. The smoking detection model 500 includes a fusion weight module 510, a pre-trained target detection model 520, a feature fusion module 530, and a pre-trained classification model 540, and the classification model 540 is a random forest algorithm.

[0119] The fusion weight module 510 is configured to determine the fusion weight of the infrared image and the visible light image according to the image display parameters when there is a face in the infrared image and / or the visible light image; the target detection model 520 is configured to determine the actual distance between the mouth of the person and the cigarette in the infrared image and the visible light image respectively; the feature fusion module 530 is configured to generate a fusion image feature according to the fusion weight and the actual distance; and the classification model 540 is configured to classify the smoking behavior of the person in the infrared image and the visible light image according to the fusion image feature to obtain a recognition result of whether there is smoking.

[0120] In an exemplary embodiment, the smoking detection model 500 can be applied to different scenes to detect the smoking behavior of personnel. For example, it can be used for smoking behavior detection of driving personnel (drivers and passengers are referred to as driving personnel) during automobile driving, for smoking behavior detection of passengers in elevators, for smoking behavior detection of personnel in offices, etc.

[0121] As shown in Figure 7 , Figure 7A processing flowchart of the smoking detection model of the present disclosure is shown. When the smoking behavior detection is performed by using the smoking detection model 500, the infrared image and the visible light image containing the same person need to be obtained, and the infrared image and the visible light image are images of the same person taken by the camera from different positions. After obtaining the infrared image and the visible light image containing the same person, it is determined whether there is a face in the infrared image and / or the visible light image by using the face recognition model. After the face recognition model determines that there is a face in the infrared image and / or the visible light image, the predicted face frame, the predicted mouth frame in the infrared image and the visible light image are also predicted, and if there is a cigarette in the image, the predicted cigarette frame of the cigarette can also be predicted. That is, if there is a face of the person in at least one of the infrared image and the visible light image, the obtained infrared image and visible light image are taken as the input of the fusion weight module 510 and the target detection model 520, and are input into the fusion weight module 510 and the target detection model 520.

[0122] The fusion weight module 520 obtains the image display parameters of the infrared image and the visible light image, and the image display parameters are image gray values. Based on the above formula (1) and formula (2), the image gray mean value and the image gray standard deviation of the infrared image are calculated according to the image gray values of the infrared image, and are represented as ir2 arg and ir2 σ respectively. The image gray mean value and the image gray standard deviation of the visible light image are calculated according to the image gray values of the visible light image, and are represented as rgb2 arg and rgb2 σ respectively. The image gray mean value and the image gray standard deviation corresponding to the infrared image and the visible light image are (ir2 arg , ir2 σ , rgb2 arg , and rgb2 σ respectively. After the normalization processing of (ir2 arg , ir2 σ , rgb2 arg , and rgb2 σ , the fusion weight (a2, b2, c2, d2) is obtained, wherein a2+b2+c2+d2=2. The fusion weight module 520 outputs the fusion weight, and inputs the fusion weight into the feature fusion module 530.

[0123] The target detection model 520 obtains the coordinates of the center point position of the predicted mouth frame and the coordinates of the center point position of the predicted cigarette frame from the infrared image, and then calculates the center point distance between the center point position of the mouth of the person and the center point position of the cigarette in the infrared image based on the above formula (3), taking the center point distance as the actual distance between the mouth of the person and the cigarette in the infrared image, denoted as e2. Similarly, the target detection model 520 obtains the coordinates of the center point position of the predicted mouth frame and the coordinates of the center point position of the predicted cigarette frame from the visible light image, and then calculates the center point distance between the center point position of the mouth of the person and the center point position of the cigarette in the visible light image based on the above formula (3), taking the center point distance as the actual distance between the mouth of the person and the cigarette in the visible light image, denoted as f2. The target detection model 520 outputs the actual distances between the mouth of the person and the cigarette in the infrared image and the visible light image respectively, and inputs the two output actual distances into the feature fusion module 530.

[0124] The feature fusion module 530 fuses the input fusion weight and the two actual distances to obtain a fused image feature, and outputs the fused image feature, denoted as W4, W4= (a2, b2, c2, d2, e2, f2), and inputs the fused image feature into the classification model 540. The classification model 540 classifies the smoking behavior of the same person in the infrared image and the visible light image according to the fused image feature, and obtains the recognition result of whether there is smoking.

[0125] Since the infrared image and the visible light image are images collected at different positions, and the infrared image contains infrared light and the visible light image contains visible light, taking the infrared image and the visible light image as the input of the smoking detection model can fully play the complementary role of the infrared image and the visible light image, and avoid the misrecognition of whether the person smokes caused by the cigarette light spot reflected on the face in a complex environment scene.

[0126] By providing the above smoking detection model, the fusion weight of the infrared image and the visible light image with respect to the image gray value, and the distance between the mouth of the person and the cigarette in the infrared image and the visible light image are taken as the fused image feature related to the smoking behavior, and after obtaining the fused image feature, the infrared image and the visible light image are fused, the complex environment scene and the smoking behavior are directly related, and the correlation between the smoking behavior and the complex environment scene can be obtained. Using the fused image feature related to the smoking behavior for smoking behavior recognition can directly observe the trend of the infrared image and the visible light image related to the smoking behavior prediction, improve the detection accuracy of whether the person smokes, and reduce the occurrence of smoking behavior misrecognition.

[0127] The following is an embodiment of a smoking behavior recognition device of the present disclosure, which can be used to implement the method embodiment of the present disclosure. For details not disclosed in the smoking behavior recognition device embodiment of the present disclosure, please refer to the method embodiment of the present disclosure.

[0128] Figure 8 A structural schematic diagram of a smoking behavior recognition device to which an embodiment of the present disclosure can be applied is shown. Please refer to Figure 8 The smoking behavior recognition device shown in the diagram can be implemented by software, hardware, or a combination of the two to become all or part of a vehicle, and can also be integrated as an independent module in the vehicle or on a server.

[0129] The smoking behavior recognition device 800 in the embodiment of the present disclosure, the smoking behavior recognition device 800 described above comprises:

[0130] The weight calculation module 810 is configured to determine the fusion weight of the person image according to the image display parameter;

[0131] The distance calculation module 820 is configured to determine the actual distance between the mouth of the person in the person image and the cigarette;

[0132] The feature fusion module 830 is configured to generate a fused image feature according to the fusion weight and the actual distance;

[0133] The behavior recognition module 840 is configured to classify the smoking behavior of the driver according to the fused image feature to obtain a recognition result of whether there is smoking.

[0134] In an exemplary embodiment, based on the foregoing scheme, the person image comprises an infrared image and a visible light image, and the smoking behavior recognition device 800 described above further comprises:

[0135] The infrared image acquisition unit is configured to acquire a first video stream of a position of the driver in the vehicle cabin, and perform frame extraction processing on the first video stream to obtain the infrared image, wherein the first video stream is collected by an infrared camera;

[0136] The visible light image acquisition unit is configured to acquire a second video stream of the overall environment inside the vehicle cabin, and perform frame extraction processing on the second video stream to obtain a vehicle cabin environment image, wherein the second video stream is collected by a visible light camera; and the image at the position is intercepted from the vehicle cabin environment image to obtain the visible light image.

[0137] In an exemplary embodiment, based on the foregoing scheme, the image display parameter comprises an image gray value, and the weight calculation module 810 described above comprises:

[0138] The first calculation unit is configured to determine a first image gray mean value and a first image gray standard deviation of the infrared image according to image gray values of the infrared image.

[0139] The second calculation unit is configured to determine a second image gray mean value and a second image gray standard deviation of the visible light image according to image gray values of the visible light image.

[0140] The fusion unit is configured to normalize the first image gray mean value, the first image gray standard deviation, the second image gray mean value and the second image gray standard deviation to obtain the fusion weight.

[0141] In an exemplary embodiment, based on the foregoing scheme, the distance calculation module 820 includes:

[0142] The position acquisition unit is configured to identify the center point position of the mouth of the person and the center point position of the cigarette in the person image by using a pre-trained target detection model.

[0143] The third calculation unit is configured to calculate a distance between the center point position of the mouth of the person and the center point position of the cigarette to obtain a center point distance.

[0144] The distance determination unit is configured to determine the center point distance as the actual distance.

[0145] In an exemplary embodiment, based on the foregoing scheme, the behavior recognition module 840 is specifically configured to input the fusion image feature into a pre-trained classification model to classify the smoking behavior of the driver to obtain the recognition result of whether there is smoking.

[0146] In an exemplary embodiment, based on the foregoing scheme, the smoking behavior recognition device 800 further includes:

[0147] The sample face image acquisition unit is configured to acquire an infrared face image and a visible light face image.

[0148] The sample fusion weight calculation unit is configured to determine a sample image gray mean value and a sample image gray standard deviation of the infrared face image and the visible light face image respectively according to image gray values to obtain a sample fusion weight.

[0149] The distance value pair calculation unit is configured to determine a distance between a mouth of a person and a cigarette in the infrared face image and the visible light face image respectively to obtain a distance value pair.

[0150] The sample image feature fusion unit is configured to generate a sample fusion image feature according to the sample fusion weight and the distance value pair.

[0151] The classification model unit is configured to train the classification model by using the sample fusion image features.

[0152] It should be noted that the smoking behavior recognition device provided in the above embodiments is used to execute the smoking behavior recognition method, and only the division of the above functional modules is used as an example. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the smoking behavior recognition device and the smoking behavior recognition method provided in the above embodiments belong to the same concept, and therefore, for details not disclosed in the device embodiments of the present disclosure, please refer to the above-mentioned smoking behavior recognition method embodiments of the present disclosure, which will not be described here.

[0153] The embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program. The program is executed by a processor to implement the steps of the method of any of the above embodiments. The computer readable storage medium can include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, micro-drives, and magneto-optical disks, ROM, RAM, EPROM, EEPROM, DRAM, VRAM, flash memory device, magnetic or optical cards, nanosystem (including molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0154] The embodiments of the present disclosure further provide a vehicle, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to implement the steps of the method of any of the above embodiments.

[0155] Figure 9 The structure of the vehicle is schematically shown. Please refer to Figure 9 As shown, the vehicle 900 includes a processor 901 and a memory 902.

[0156] In the embodiments of the present disclosure, the processor 901 is the control center of the computer system, which can be a processor of a physical machine or a processor of a virtual machine. The processor 901 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 901 can be implemented in at least one of the hardware forms of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 901 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state.

[0157] In the embodiments of the present disclosure, the processor 901 is specifically configured to: determine a fusion weight of a person image according to an image display parameter; the person image is an image of a position where a driver in a vehicle cabin is located; determine an actual distance between a mouth of a person in the person image and a cigarette; generate a fusion image feature according to the fusion weight and the actual distance; and classify a smoking behavior of the driver according to the fusion image feature to obtain an identification result of whether there is smoking.

[0158] Further, the person image includes an infrared image and a visible light image, and the processor 901 is further configured to: acquire a first video stream of the position where the driver in the vehicle cabin is located, and perform frame extraction processing on the first video stream to obtain the infrared image, the first video stream being collected by an infrared camera; acquire a second video stream of an overall environment inside the vehicle cabin, and perform frame extraction processing on the second video stream to obtain a vehicle cabin environment image, the second video stream being collected by a visible light camera; and obtain the visible light image by intercepting an image at the position from the vehicle cabin environment image.

[0159] Further, the image display parameter includes an image gray value, and the processor 901 is further configured to: determine a first image gray average value and a first image gray standard deviation of the infrared image according to an image gray value of the infrared image; determine a second image gray average value and a second image gray standard deviation of the visible light image according to an image gray value of the visible light image; and perform normalization processing on the first image gray average value, the first image gray standard deviation, the second image gray average value, and the second image gray standard deviation to obtain the fusion weight.

[0160] Further, the processor 901 is further configured to: identify the center point position of the mouth of the person and the center point position of the cigarette in the person image by using a pre-trained object detection model; calculate a distance between the center point position of the mouth of the person and the center point position of the cigarette, to obtain a center point distance; and determine the center point distance as the actual distance.

[0161] Further, the processor 901 is further configured to: input the fused image feature into a pre-trained classification model, to classify the smoking behavior of the driver, to obtain the identification result of whether there is smoking.

[0162] Further, the processor 901 is further configured to: obtain an infrared face image and a visible light face image; determine a sample image gray average value and a sample image gray standard deviation of each of the infrared face image and the visible light face image according to image gray values, to obtain a sample fusion weight; determine a distance between the mouth of a person and a cigarette in each of the infrared face image and the visible light face image, to obtain a distance value pair; generate a sample fused image feature according to the sample fusion weight and the distance value pair; and train the classification model by using the sample fused image feature.

[0163] The memory 902 can include one or more computer readable storage media. The computer readable storage media can be non-transitory. The memory 902 can also include high-speed random access memory and non-volatile memory such as one or more magnetic disk storage devices, flash memory devices. In some embodiments of the present disclosure, the non-transitory computer readable storage medium in the memory 902 is used to store at least one instruction for being executed by the processor 901 to implement the method in the embodiments of the present disclosure.

[0164] In some embodiments, the vehicle 900 further includes a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902 and the peripheral device interface 903 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 903 through a bus, a signal line or a circuit board. Specifically, the peripheral device includes at least one of a display screen 904, a camera 905 and an audio circuit 906.

[0165] The peripheral interface 903 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 901 and the memory 902. In some embodiments of the present disclosure, the processor 901, the memory 902 and the peripheral interface 903 are integrated on the same chip or circuit board; in some other embodiments of the present disclosure, any one or two of the processor 901, the memory 902 and the peripheral interface 903 can be implemented on a separate chip or circuit board. The embodiments of the present disclosure do not make specific limitations in this regard.

[0166] The display screen 904 is used to display a UI (User Interface). The UI can include graphics, text, icons, videos and any combination thereof. When the display screen 904 is a touch display screen, the display screen 904 also has the ability to collect touch signals on or above the surface of the display screen 904. The touch signals can be input as control signals to the processor 901 for processing. At this time, the display screen 904 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments of the present disclosure, the display screen 904 can be one, arranged on the front panel of the vehicle 900; in some other embodiments of the present disclosure, the display screen 904 can be at least two, arranged on different surfaces of the vehicle 900 or in a folding design; in some other embodiments of the present disclosure, the display screen 904 can be a flexible display screen, arranged on a curved surface or a folding surface of the vehicle 900. Even, the display screen 904 can also be arranged in an irregular shape other than a rectangle, that is, a special-shaped screen. The display screen 904 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0167] The camera 905 is used to collect images or videos. Optionally, the camera 905 includes a front camera and a rear camera. Generally, the front camera is arranged on the front panel of the vehicle 900, and the rear camera is arranged on the back of the vehicle 900. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera and a telephoto camera, to realize the background blur function of the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function of the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments of the present disclosure, the camera 905 can also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0168] The audio circuit 906 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into an electrical signal input to the processor 901 for processing. The microphone can be multiple for the purpose of stereo sound collection or noise reduction, and arranged at different parts of the vehicle 900 respectively. The microphone can also be an array microphone or an omnidirectional collection microphone.

[0169] The power supply 907 is used to supply power to various components in the vehicle 900. The power supply 907 can be alternating current, direct current, disposable battery or rechargeable battery. When the power supply 907 includes a rechargeable battery, the rechargeable battery can be a wired charging battery or a wireless charging battery. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0170] The structural block diagram of the vehicle 900 shown in the embodiments of the present disclosure does not constitute a limitation on the vehicle 900, and the vehicle 900 can include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0171] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in relevant countries and regions. For example, the object features, interaction behavior features and user information involved in the present specification are obtained under full authorization.

[0172] In the description of the present disclosure, it should be understood that the terms "first", "second" and the like are only used for descriptive purpose and cannot be understood as indicating or implying relative importance. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to the specific circumstances. In addition, in the description of the present disclosure, "multiple" means two or more, unless otherwise specified. "And / or", which describes the association between the associated objects, means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship.

[0173] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, equivalent changes made according to the claims of the present disclosure are still within the scope covered by the present disclosure.

Claims

1. A smoking behavior recognition method, characterized by, The smoking behavior recognition method comprises: obtaining an image of a position of a driver in a vehicle cabin to obtain a person image, the person image comprising an infrared image and a visible light image; determining a first image gray average value and a first image gray standard deviation of the infrared image according to an image gray value of the infrared image; determining a second image gray average value and a second image gray standard deviation of the visible light image according to an image gray value of the visible light image; normalizing the first image gray average value, the first image gray standard deviation, the second image gray average value and the second image gray standard deviation to obtain a fusion weight of the person image; determining an actual distance between a mouth of a person in the person image and a cigarette; fusing the fusion weight and the actual distance to obtain a fusion image feature; classifying a smoking behavior of the driver according to the fusion image feature to obtain a recognition result of whether there is smoking.

2. The smoking behavior recognition method of claim 1, wherein, The obtaining of the image of the position of the driver in the vehicle cabin to obtain the person image comprises: obtaining a first video stream of the position of the driver in the vehicle cabin and performing frame extraction processing on the first video stream to obtain the infrared image, the first video stream being collected by an infrared camera; obtaining a second video stream of an overall environment inside the vehicle cabin and performing frame extraction processing on the second video stream to obtain a vehicle cabin environment image, the second video stream being collected by a visible light camera; extracting an image at the position from the vehicle cabin environment image to obtain the visible light image.

3. The smoking behavior recognition method of claim 1, wherein, The determination of the actual distance between the mouth of the person in the person image and the cigarette comprises: recognizing a center point position of the mouth of the person and a center point position of the cigarette in the person image by using a pre-trained object detection model; calculating a distance between the center point position of the mouth of the person and the center point position of the cigarette to obtain a center point distance; determining the center point distance as the actual distance.

4. The smoking behavior recognition method according to any one of claims 1 to 3, characterized by, The classification of the smoking behavior of the driver according to the fusion image feature to obtain the recognition result of whether there is smoking comprises: inputting the fusion image feature into a pre-trained classification model to classify the smoking behavior of the driver to obtain the recognition result of whether there is smoking.

5. The smoking behavior recognition method of claim 4, wherein, Before the classification of the smoking behavior of the driver according to the fusion image feature to obtain the recognition result of whether there is smoking, the smoking behavior recognition method further comprises: obtaining an infrared face image and a visible light face image; determining a sample image gray average value and a sample image gray standard deviation of the infrared face image and the visible light face image respectively according to image gray values to obtain a sample fusion weight; determining a distance between a mouth of a person in the infrared face image and the visible light face image respectively to obtain a distance value pair; generating a sample fusion image feature according to the sample fusion weight and the distance value pair; training the classification model by using the sample fusion image feature.

6. A smoking detection model, characterized in that, The smoking detection model comprises: a fusion weight module configured to, in a case where a human face exists in an infrared image and / or a visible light image, calculate an image gray average value and an image gray standard deviation of the infrared image according to image gray values of the infrared image, calculate an image gray average value and an image gray standard deviation of the visible light image according to image gray values of the visible light image, and normalize the image gray average value and the image gray standard deviation of the infrared image and the visible light image respectively to obtain fusion weights of the infrared image and the visible light image; a pre-trained target detection model configured to determine an actual distance between a human mouth and a cigarette in the infrared image and the visible light image respectively; a feature fusion module configured to fuse the fusion weights and the actual distance to obtain a fused image feature; a pre-trained classification model configured to classify a smoking behavior of a human in the infrared image and the visible light image according to the fused image feature to obtain a recognition result of whether smoking exists.

7. A smoking behavior recognition apparatus, characterized by, The smoking behavior recognition device comprises: a weight calculation module configured to obtain an image of a position where a driver is located in a vehicle cabin to obtain a human image, the human image comprising an infrared image and a visible light image; determine a first image gray average value and a first image gray standard deviation of the infrared image according to image gray values of the infrared image; determine a second image gray average value and a second image gray standard deviation of the visible light image according to image gray values of the visible light image; and normalize the first image gray average value, the first image gray standard deviation, the second image gray average value and the second image gray standard deviation to obtain fusion weights of the human image; a distance calculation module configured to determine an actual distance between a human mouth and a cigarette in the human image; a feature fusion module configured to fuse the fusion weights and the actual distance to obtain a fused image feature; a behavior recognition module configured to classify a smoking behavior of the driver according to the fused image feature to obtain a recognition result of whether smoking exists.

8. A vehicle comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the smoking behavior recognition method according to any one of claims 1 to 5.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the smoking behavior recognition method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Occupant monitoring device and occupant monitoring method

    WO2023105751A1