Driver and passenger behavior recognition method, device, vehicle and storage medium
By using a combination of multi-angle camera equipment and attention mechanism layer, the problem of low behavior recognition accuracy caused by the occlusion of the driver or occupant is solved, and a higher accuracy of the driver and passenger behavior recognition is achieved.
Patent Information
- Application Number
- CN202210044622.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-01-14
AI Technical Summary
In the prior art, since the camera equipment is placed in a fixed position and the space in the cockpit is small, the driver or occupant is often blocked or exceeded the shooting range of the camera equipment, affecting the accuracy of the recognition of the driver or occupant's behavior.
At least two imaging devices are used to capture images from different angles, and the preferred images with a degree of occlusion lower than a preset degree are screened through the occlusion recognition model, and then the weight of image information that has an effect on behavior recognition is increased using a behavior recognition model including an attention mechanism layer.
The accuracy of the recognition of the behavior of drivers and passengers has been improved, and the accuracy of behavior recognition has been further improved through preliminary occlusion analysis and the processing of attention mechanism.
Smart Images

Figure CN115116040B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a method, device, vehicle and storage medium for identifying the behaviors of vehicle occupants. Background Art
[0002] During the driving process of a vehicle, the driver may have illegal driving behaviors, or the passengers may have illegal interference behaviors, leading to dangers. For example, the driver has bad driving behaviors such as answering or making calls, driving while fatigued, or searching for items during driving; or the passengers have illegal behaviors of snatching the steering wheel.
[0003] Currently, it mainly relies on in-vehicle camera devices to collect images containing the driver and / or passengers, and uses algorithms to monitor and analyze the behaviors of the driver or passengers in the images. However, due to the fixed placement of the camera devices and the small space inside the cockpit, the driver or passengers are often blocked or out of the shooting range of the camera devices, thus affecting the accuracy of identifying the behaviors of the driver or passengers. Summary of the Invention
[0004] Embodiments of this application provide a method, device, vehicle and storage medium for identifying the behaviors of vehicle occupants, which can solve the problem of low accuracy in identifying the behaviors of the driver or passengers.
[0005] In a first aspect, embodiments of this application provide a method for identifying the behaviors of vehicle occupants, and the method includes:[[]]
[0006] Obtain to-be-identified images respectively captured by at least two camera devices; the at least two camera devices respectively have different shooting angles;
[0007] Input each to-be-identified image into an occlusion recognition model to obtain a preferred to-be-identified image; the preferred to-be-identified image is an image with the occlusion degree of the vehicle occupant lower than a preset degree;
[0008] Input the preferred to-be-identified image into a behavior recognition model including an attention mechanism layer for processing to obtain a behavior recognition result of the vehicle occupant; the attention mechanism layer is used to increase the weight of the image information in the preferred to-be-identified image that has an impact on behavior recognition when processing the preferred to-be-identified image.
[0009] In a second aspect, embodiments of this application provide a device for identifying the behaviors of vehicle occupants, and the device includes:[[]]
[0010] An obtaining module, configured to obtain to-be-identified images respectively captured by at least two camera devices; the at least two camera devices respectively have different shooting angles;
[0011] An input module for inputting an image to be recognized into an occlusion recognition model to obtain a preferred image to be recognized; the preferred image to be recognized is an image in which the occlusion degree of the driver and passengers is lower than a preset degree.
[0012] A generation module for inputting the preferred image to be recognized into a behavior recognition model including an attention mechanism layer for processing to obtain a behavior recognition result of the driver and passengers; the attention mechanism layer is used to increase the weight of the image information in the preferred image to be recognized that has an impact on behavior recognition when processing the preferred image to be recognized.
[0013] In a third aspect, an embodiment of the present application provides a vehicle, including a terminal device. The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of the first aspect as described above is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method of the first aspect as described above is implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is caused to execute the method of the first aspect as described above.
[0016] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The in-vehicle device inputs the images to be recognized taken from at least two different shooting angles into the occlusion recognition model to initially analyze the occlusion degree. In this way, at least one preferred image to be recognized with an occlusion degree lower than the preset degree is determined from at least two images to be recognized, so as to initially improve the accuracy of recognizing the behavior of the driver and passengers. Then, the preferred image to be recognized is input into the behavior recognition model including the attention mechanism layer for processing, so as to use the attention mechanism layer to increase the weight of the image information in the preferred image to be recognized that has an impact on behavior recognition, so as to further improve the accuracy of recognizing the behavior of the driver and passengers. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of the implementation of a method for recognizing the behavior of a driver and passengers provided by an embodiment of the present application;
[0019] Figure 2 It is a flowchart of the implementation of a method for identifying the behavior of vehicle occupants provided by another embodiment of the present application;
[0020] Figure 3 It is a schematic diagram of an implementation manner of S103 of a method for identifying the behavior of vehicle occupants provided by an embodiment of the present application;
[0021] Figure 4 It is a schematic diagram of the structure of a behavior recognition model in a method for identifying the behavior of vehicle occupants provided by an embodiment of the present application;
[0022] Figure 5 It is a schematic diagram of the structure of a device for identifying the behavior of vehicle occupants provided by an embodiment of the present application;
[0023] Figure 6 It is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0024] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0025] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0026] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0027] The method for identifying the behavior of vehicle occupants provided by the embodiments of the present application can be applied to terminal devices such as tablet computers, vehicle-mounted devices, laptop computers, ultra-mobile personal computers (UMPCs), etc. The embodiments of the present application do not impose any restrictions on the specific types of terminal devices.
[0028] In this embodiment, the above-mentioned method for identifying the behavior of vehicle occupants can be specifically applied to in-vehicle devices. Among them, at least two camera devices are installed in the in-vehicle device to respectively capture the driver's seat in the intelligent vehicle. Each camera device usually has a different shooting angle. In addition, the above-mentioned vehicle occupants include the driver and passengers. Since passengers may have illegal behaviors such as snatching the steering wheel or hitting the driver, when capturing the driver's seat, the captured images to be identified may include both the driver and passengers at the same time.
[0029] Please refer to Figure 1 , Figure 1 which shows a flowchart for implementing a method for identifying the behavior of vehicle occupants provided by an embodiment of the present application. The method includes the following steps:
[0030] S101. The in-vehicle device obtains images to be identified respectively captured by at least two camera devices; the at least two camera devices respectively have different shooting angles.
[0031] In one embodiment, the above-mentioned in-vehicle device is communicatively connected to the camera device, and can obtain the images to be identified from each camera device in real time or at preset intervals. In this embodiment, the following description is made by taking the number of camera devices as two as an example.
[0032] Among them, the two camera devices can be pre-classified into a main camera device and a secondary camera device. Usually, the two camera devices are respectively fixed at different positions to capture the driver's seat. In addition, since each camera device is fixed and the space in the cockpit is small, the driver or passengers in the captured images to be identified are often blocked or outside the shooting range of the two camera devices, thus affecting the judgment of the in-vehicle device on the behavior recognition result of the vehicle occupants.
[0033] S102. The in-vehicle device inputs each image to be identified into an occlusion recognition model to obtain a preferred image to be identified; the preferred image to be identified is an image in which the occlusion degree of the vehicle occupant is lower than a preset degree.
[0034] In one embodiment, the above-mentioned occlusion recognition model can be an existing neural network model, which can process each image to be identified and output a preferred image to be identified in which the occlusion degree of the vehicle occupant is lower than a preset degree.
[0035] In one embodiment, the above preset degree can be set by the staff according to the actual situation, and there is no limitation in this regard. It should be noted that if the occlusion recognition model processes each image to be recognized and determines that the occlusion degree of the driver and passengers is lower than the preset degree, the final classification result is that each image to be recognized is an optimal image to be recognized. Or, if the occlusion recognition model processes each image to be recognized and determines that the occlusion degree of the driver and passengers is not lower than the preset degree, the final classification result can also be that each image to be recognized is an optimal image to be recognized. That is, although the occlusion degree of the driver and passengers in each image to be recognized is not lower than the preset degree, during the subsequent processing, the behavior recognition model can predict by combining the behaviors of the unoccluded driver and passengers in each image to be recognized.
[0036] In a specific embodiment, when there are two camera devices, the image to be recognized captured by the main camera device is the main captured image to be recognized, and the image to be recognized captured by the secondary camera device is the secondary captured image to be recognized. Among them, the occlusion recognition model is specifically a ResNet50 neural network model. The classification of the last fully connected layer in the ResNet50 neural network model is a three-class classification. Specifically, the three classifications are: the main captured image to be recognized is an optimal image to be recognized, the secondary captured image to be recognized is an optimal image to be recognized, and both images to be recognized are optimal images to be recognized, and there is no limitation in this regard.
[0037] In one embodiment, before inputting both the main captured image to be recognized and the secondary captured image to be recognized into the occlusion recognition model, the vehicle-mounted device can merge the two images to be recognized to generate a merged image to be recognized. When the occlusion recognition model processes the merged image to be recognized, it can compare the two images to be recognized simultaneously to further determine the optimal image to be recognized with a lower occlusion degree. Furthermore, when the classification result is that both images to be recognized are optimal images to be recognized, the vehicle-mounted device can focus on increasing the weight of the optimal image to be recognized with a lower occlusion degree through the behavior recognition model to improve the accuracy of the behavior recognition result.
[0038] Among them, the merged image to be recognized will contain the image features of the main captured image to be recognized and the secondary captured image to be recognized. Among them, in the merged image to be recognized, the main captured image to be recognized and the secondary captured image to be recognized should be of the same size.
[0039] S103. The vehicle-mounted device inputs the optimal image to be recognized into the behavior recognition model including an attention mechanism layer for processing to obtain the behavior recognition result of the driver and passengers; the attention mechanism layer is used to increase the weight of the image information in the optimal image to be recognized that has an impact on behavior recognition when processing the optimal image to be recognized.
[0040] In one embodiment, when processing the preferably to-be-recognized image, the above-mentioned attention mechanism layer can increase the weight of the image information in the preferably to-be-recognized image that has an impact on behavior recognition. Specifically, the above-mentioned attention mechanism can be used to calculate the weights in the image features of the preferably to-be-recognized image that are respectively helpful for behavior recognition, and can obtain more detailed information that needs to be concerned from the image features of the preferably to-be-recognized image according to the weights, and suppress other useless information in the preferably to-be-recognized image. Furthermore, the image features obtained after being processed by the attention mechanism can represent the important features in the to-be-recognized image that have an impact on behavior recognition, and improve the judgment accuracy of the behavior recognition result.
[0041] In one embodiment, the above-mentioned behavior recognition results include but are not limited to: bad driving behaviors such as a driver making a call, driving while fatigued, or looking for an item; or a passenger's illegal behavior of snatching the steering wheel, which is not limited here.
[0042] In this embodiment, the in-vehicle device inputs the to-be-recognized images taken from at least two different shooting angles into the occlusion recognition model to preliminarily analyze the occlusion degree. In this way, at least one preferably to-be-recognized image with an occlusion degree lower than a preset degree is determined from the at least two to-be-recognized images to preliminarily improve the accuracy of recognizing the behaviors of the driver and passengers. Then, the preferably to-be-recognized image is input into the behavior recognition model including the attention mechanism layer for processing, so as to use the attention mechanism layer to increase the weight of the image information in the preferably to-be-recognized image that has an impact on behavior recognition, and further improve the accuracy of recognizing the behaviors of the driver and passengers.
[0043] In one embodiment, when the behavior recognition model processes the input data, the dimension of the input data should be consistent with the dimension that the behavior recognition model can process. However, the dimension that the behavior recognition model can process is usually fixed, but the number of preferably to-be-recognized images may be one or multiple. Therefore, in order to enable the behavior recognition model to process any number of preferably to-be-recognized images and ensure the prediction accuracy of the behavior recognition result, the in-vehicle device also needs to preprocess any number of preferably to-be-recognized images. Specifically, referring to Figure 2 , the in-vehicle device can process any number of preferably to-be-recognized images according to the following steps S11 - S14:
[0044] S11. If there is one preferably to-be-recognized image, the in-vehicle device copies the preferably to-be-recognized image to obtain multiple copied preferably to-be-recognized images.
[0045] S12. If there are multiple preferably to-be-recognized images, the in-vehicle device keeps the multiple to-be-recognized images unchanged.
[0046] In one embodiment, the above-mentioned preferred image to be recognized is one, that is, in step S102, the final classification result is that the main camera image to be recognized or the secondary camera image to be recognized is the preferred image to be recognized. Also, when there are multiple preferred images to be recognized, it means that the final classification result is that both of the two images to be recognized are preferred images to be recognized.
[0047] It should be noted that when there is one preferred image to be recognized, the vehicle-mounted device can copy the preferred image to be recognized once to obtain two copied preferred images to be recognized. Also, when there are two preferred images to be recognized, the multiple images to be recognized remain unchanged.
[0048] Among them, the purpose of copying the preferred image to be recognized is as follows: compared with only inputting one preferred image to be recognized into the behavior recognition model for processing, when the multiple copied preferred images to be recognized are processed through the attention mechanism layer in the behavior recognition model, the weights of the image information in each preferred image to be recognized that affects behavior recognition can be increased respectively, so as to further enhance more image detail information that needs to be concerned in the image to be recognized and suppress other useless information.
[0049] It should be added that another purpose of copying the preferred image to be recognized is as follows: on the premise that the number of preferred images to be recognized is uncertain, the vehicle-mounted device can perform a copying process on the preferred image to be recognized and splice the multiple copied preferred images to be recognized, so that the dimension of the spliced preferred image to be recognized is consistent with the dimension of the image that the behavior recognition model can process. In this way, the dimension of the preferred image to be recognized input into the behavior recognition model can be unified. Specifically:
[0050] S13. The vehicle-mounted device respectively determines the channel dimension of each preferred image to be recognized.
[0051] S14. The vehicle-mounted device splices each preferred image to be recognized in the channel dimension to obtain the spliced preferred image to be recognized; the channel dimension of the spliced preferred image to be recognized is a preset dimension.
[0052] In one embodiment, the above-mentioned channel dimension is the number of channels of the preferred image to be recognized. Generally, the channel dimension of one preferred image to be recognized is 3, namely the three channels of Red, Green, and Blue, that is, RGB.
[0053] In one embodiment, the above-mentioned preset dimension can be set by the staff according to the actual situation, and its dimension can be a multiple of 3, which is not limited herein. In this embodiment, the above-mentioned preset dimension can be 6. Based on this, when there is 1 preferred image to be recognized, it only needs to be copied once and then spliced. When there are 2 preferred images to be recognized, they only need to be directly spliced, which is not limited herein.
[0054] Exemplarily, if the dimension of a single preferred image to be recognized is 640*480*3, then after replication and splicing, the dimension of the generated preferred image to be recognized will become 640*480*6. Among them, 640*480 is the size of the preferred image to be recognized, 3 is the channel dimension of the preferred image to be recognized, and 6 is the preset dimension.
[0055] It can be understood that when there are more than two preferred images to be recognized and the dimension after directly splicing the more than two preferred images to be recognized is not the preset dimension, then the multiple preferred images to be recognized need to be processed separately. Specifically, when the dimension after directly splicing the multiple preferred images to be recognized is lower than the preset dimension, replication is also required to increase the dimension of the preferred image to be recognized after splicing; and when the dimension after directly splicing the multiple preferred images to be recognized is higher than the preset dimension, the lowest preferred image to be recognized can be deleted sequentially according to the occlusion degree of each preferred image to be recognized until the dimension after direct splicing is equal to the preset dimension.
[0056] In one embodiment, the above-mentioned behavior recognition model includes at least one model processing node, and each model processing node includes an attention mechanism layer and at least one convolutional layer; referring to Figure 3 , the behavior recognition model can process the spliced preferred image to be recognized through the following steps S1-S5 to obtain the behavior recognition result:
[0057] S1. The vehicle-mounted device uses the image features of the preferred image to be recognized as the current input features.
[0058] S2. The vehicle-mounted device inputs the current input features into the attention mechanism layer and the convolutional layer in the current model processing node respectively for processing to obtain attention features and convolutional features.
[0059] S3. The vehicle-mounted device splices the attention features and the convolutional features to obtain the spliced current input features.
[0060] In one embodiment, the above-mentioned behavior recognition model includes multiple model processing nodes, and each model processing node includes an attention mechanism layer and three convolutional layers.
[0061] Specifically, referring to Figure 4 , Figure 4The Block in it is a model processing node, and there are N in total. Among them, each model processing node sequentially includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The convolutional strides of the first convolutional layer and the third convolutional layer are both 1, and the convolutional stride of the second convolutional layer is 3. In addition, Attention is an attention mechanism layer, and specific reference can be made to Figure 4 the network layer composed of Global pooling, FC, FC, and sigmoid in it. Among them, Global pooling is a global pooling layer, the first FC below it is a dimension compression layer, the FC below that is a dimension reconstruction layer, and sigmoid is an activation layer.
[0062] Specifically, inputting the current input feature into the attention mechanism layer in the current model processing node means processing the current input feature sequentially through the above-mentioned Global pooling, two FCs, and sigmoid to obtain an attention feature. Inputting the current input feature into the convolutional layer in the current model processing node for processing means processing it sequentially through the first convolutional layer, the second convolutional layer, and the third convolutional layer to obtain a convolutional feature.
[0063] It should be added that the dimensions of the attention feature and the convolutional feature obtained through the above processing are the same. Therefore, the attention feature and the convolutional feature can be directly concatenated to obtain the concatenated current input feature.
[0064] It should be noted that for the global pooling layer in the attention mechanism layer, it can be a max pooling layer. Among them, the max pooling layer can select the features with the largest values in the local feature maps from the feature maps of the preferred image to be recognized (the concatenated preferred image to be recognized is an image with 6 channel dimensions, and the current input feature of the preferred image to be recognized consists of 6 feature maps) to form a new feature map, so as to reduce the error of the behavior recognition model in extracting features from the preferred image to be recognized.
[0065] Among them, the first FC is used to compress the dimension of the current input feature to aggregate the spatial information between each feature map. Then, it is input into the second FC, which is used to reconstruct the dimension of the compressed current input feature so that the reconstructed dimension is the same as the dimension of the current input feature. Based on this, the current input feature reconstructed according to the aggregated spatial information between each feature map can further represent more detailed information in the preferred image to be recognized. Among them, inputting the reconstructed current input feature into the activation function for processing is mainly to retain the main feature parameters in the current input feature.
[0066] Based on the above description, after processing the preferred image to be recognized using the attention mechanism layer, enhancement processing such as color difference, sharpening, and brightness can be performed on the preferred image to be recognized, so that the attention features obtained after processing can generate new current input features that are more conducive to expressing the preferred image to be recognized after being concatenated with the convolutional features.
[0067] S4. If the current model processing node is the terminal processing node, the vehicle-mounted device outputs a behavior recognition result based on the current input features.
[0068] S5. If the current model processing node is a non-terminal processing node, the vehicle-mounted device determines the next processing node of the current model processing node as the new current model processing node.
[0069] It should be noted that after executing step S5, the vehicle-mounted device should input the concatenated current input features into the new current model processing node for processing again. That is, repeat steps S2 - S5 until a behavior recognition result is obtained.
[0070] It can be understood that when there are multiple model processing nodes, the attention features and convolutional features output by each model processing node need to be concatenated, and the concatenated features are used as the current input features of the next model processing node for processing until a behavior recognition result is obtained. At this time, the accuracy of the behavior recognition result generated by processing the current input features according to the above attention mechanism will be relatively high.
[0071] Please refer to Figure 5 , Figure 5 which is a structural block diagram of a device for recognizing the behavior of vehicle occupants provided in an embodiment of the present application. In this embodiment, each module included in the device for recognizing the behavior of vehicle occupants is used to execute Figures 1 to 4 the respective steps in the corresponding embodiment. Specifically, please refer to Figures 1 to 4 and Figures 1 to 4 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 5 , the device 500 for recognizing the behavior of vehicle occupants may include: an acquisition module 510, an input module 520, and a generation module 530, where:
[0072] The acquisition module 510 is configured to acquire images to be recognized captured by at least two camera devices; the at least two camera devices have different shooting angles.
[0073] The input module 520 is configured to input the images to be recognized into an occlusion recognition model to obtain preferred images to be recognized; the preferred images to be recognized are images in which the occlusion degree of the vehicle occupants is lower than a preset degree.
[0074] A generation module 530 is configured to input a preferably to-be-recognized image into a behavior recognition model including an attention mechanism layer for processing, so as to obtain a behavior recognition result of the driver and passenger; the attention mechanism layer is configured to increase the weight of the image information in the preferably to-be-recognized image that has an impact on behavior recognition when processing the preferably to-be-recognized image.
[0075] In one embodiment, the driver and passenger behavior recognition device 500 further includes:
[0076] A merging module is configured to merge each to-be-recognized image respectively to obtain a merged to-be-recognized image.
[0077] In one embodiment, the driver and passenger behavior recognition device 500 further includes:
[0078] A copying module is configured to, if the preferably to-be-recognized image is one, copy the preferably to-be-recognized image to obtain multiple copied preferably to-be-recognized images;
[0079] A maintaining module is configured to, if the preferably to-be-recognized images are multiple, keep the multiple to-be-recognized images unchanged.
[0080] In one embodiment, the driver and passenger behavior recognition device 500 further includes:
[0081] A splicing module is configured to splice multiple preferably to-be-recognized images to obtain a spliced preferably to-be-recognized image.
[0082] In one embodiment, the splicing module is further configured to:
[0083] respectively determine the channel dimension of each preferably to-be-recognized image; splice each preferably to-be-recognized image in the channel dimension to obtain a spliced preferably to-be-recognized image; the channel dimension of the spliced preferably to-be-recognized image is a preset dimension.
[0084] In one embodiment, the behavior recognition model includes at least one model processing node, and each model processing node includes an attention mechanism layer and at least one convolutional layer; the generation module 530 is further configured to:
[0085] S1. Use the image features of the preferably to-be-recognized image as the current input features; S2. Input the current input features into the attention mechanism layer and the convolutional layer in the current model processing node respectively for processing to obtain attention features and convolutional features; S3. Concatenate the attention features and the convolutional features to obtain the concatenated current input features; S4. If the current model processing node is the terminal processing node, output the behavior recognition result according to the current input features; S5. If the current model processing node is a non-terminal processing node, determine the next processing node of the current model processing node as the new current model processing node; S6. Repeat steps S2 - S5 for the concatenated current input features until the behavior recognition result is obtained.
[0086] In one embodiment, each model processing node sequentially includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The convolutional strides of the first convolutional layer and the third convolutional layer are 1, and the convolutional stride of the second convolutional layer is 3; the attention mechanism layer is composed of a global pooling layer, a dimension compression layer, a dimension reconstruction layer, and an activation layer.
[0087] It should be understood that Figure 5 in the structural block diagram of the occupant behavior recognition device shown, each module is used to execute Figures 1 to 4 the respective steps in the corresponding embodiments, and for Figures 1 to 4 the respective steps in the corresponding embodiments have been explained in detail in the above embodiments. For details, please refer to Figures 1 to 4 and Figures 1 to 4 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.
[0088] Figure 6 is the structural block diagram of a terminal device provided in an embodiment of the present application. As Figure 6 shown, the terminal device 600 in this embodiment includes: a processor 610, a memory 620, and a computer program 630 stored in the memory 620 and executable on the processor 610, such as the program of the occupant behavior recognition method. When the processor 610 executes the computer program 630, it implements the steps in each of the above embodiments of the occupant behavior recognition method, such as Figure 1 S101 to S103 shown. Alternatively, when the processor 610 executes the computer program 630, it implements the functions of each module in the above Figure 5 corresponding embodiments, for example, Figure 5 the functions of the modules 510 to 530 shown. For details, please refer to Figure 5 the relevant descriptions in the corresponding embodiments.
[0089] Exemplarily, the computer program 630 can be divided into one or more modules. One or more modules are stored in the memory 620 and executed by the processor 610 to implement the occupant behavior recognition method provided in the embodiments of the present application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 630 in the terminal device 600. For example, the computer program 630 can implement the occupant behavior recognition method provided in the embodiments of the present application.
[0090] The terminal device 600 may include, but is not limited to, a processor 610 and a memory 620. Those skilled in the art can understand that Figure 6 merely examples of the terminal device 600, which do not constitute a limitation on the terminal device 600. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may also include input and output devices, network access devices, buses, etc.
[0091] The so-called processor 610 may be a central processing unit, or may also be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0092] The memory 620 may be an internal storage unit of the terminal device 600, such as the hard disk or memory of the terminal device 600. The memory 620 may also be an external storage device of the terminal device 600, such as a plug-in hard disk, a smart memory card, a flash memory card, etc. equipped on the terminal device 600. Further, the memory 620 may also include both the internal storage unit and the external storage device of the terminal device 600.
[0093] The embodiments of the present application provide a vehicle, including a terminal device. The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the occupant behavior recognition method in the above-mentioned various embodiments.
[0094] The embodiments of the present application provide a computer-readable storage medium, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the occupant behavior recognition method in the above-mentioned various embodiments.
[0095] The embodiments of the present application provide a computer program product. When the computer program product runs on the terminal device, it causes the terminal device to execute the occupant behavior recognition method in the above-mentioned various embodiments.
[0096] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for identifying the behavior of vehicle occupants, characterized in that, The method includes: Obtaining to-be-recognized images respectively captured by at least two camera devices; the at least two camera devices respectively have different shooting angles; Inputting each of the to-be-recognized images into an occlusion recognition model to obtain a preferred to-be-recognized image; the preferred to-be-recognized image is an image in which the occlusion degree of the driver and passengers is lower than a preset degree; Inputting the preferred to-be-recognized image into a behavior recognition model including an attention mechanism layer for processing to obtain a behavior recognition result of the driver and passengers; the attention mechanism layer is used to increase the weight of the image information in the preferred to-be-recognized image that has an impact on behavior recognition when processing the preferred to-be-recognized image; Wherein, the method further includes: Stitching a plurality of the preferred to-be-recognized images to obtain the stitched preferred to-be-recognized image; The stitching of the plurality of the preferred to-be-recognized images includes: Respectively determining the channel dimension of each of the preferred to-be-recognized images; Stitching each of the preferred to-be-recognized images in the channel dimension to obtain the stitched preferred to-be-recognized image; the channel dimension of the stitched preferred to-be-recognized image is a preset dimension, and the preset dimension is a multiple of 3.
2. The method according to claim 1, wherein Before inputting each of the to-be-recognized images into the occlusion recognition model to obtain a preferred to-be-recognized image, it further includes: Respectively merging each of the to-be-recognized images to obtain the merged to-be-recognized image.
3. The method according to claim 1 or 2, characterized in that, Before inputting the preferred to-be-recognized image into the behavior recognition model including an attention mechanism layer, it further includes: If there is one preferred to-be-recognized image, copying the preferred to-be-recognized image to obtain a plurality of copied preferred to-be-recognized images; If there are multiple preferred to-be-recognized images, keeping the multiple to-be-recognized images unchanged.
4. The method according to claim 1 or 2, characterized in that, The behavior recognition model includes at least one model processing node, and each model processing node includes an attention mechanism layer and at least one convolutional layer; The inputting the preferred to-be-recognized image into the behavior recognition model including an attention mechanism layer for processing to obtain a behavior recognition result of the driver and passengers includes: S1. Using the image features of the preferred to-be-recognized image as the current input features; S2. Respectively inputting the current input features into the attention mechanism layer and the convolutional layer in the current model processing node for processing to obtain attention features and convolutional features; S3. Stitching the attention features and the convolutional features to obtain the stitched current input features; S4. If the current model processing node is the terminal processing node, outputting the behavior recognition result according to the current input features; S5. If the current model processing node is a non-terminal processing node, determining the next processing node of the current model processing node as the new current model processing node; S6. Repeating steps S2 - S5 for the stitched current input features until the behavior recognition result is obtained.
5. The method according to claim 4, characterized in that Each of the model processing nodes sequentially includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The convolutional strides of the first convolutional layer and the third convolutional layer are both 1, and the convolutional stride of the second convolutional layer is 3. The attention mechanism layer is composed of a global pooling layer, a dimension compression layer, a dimension reconstruction layer, and an activation layer.
6. An occupant behavior recognition device, characterized in that, The apparatus includes: an acquisition module, configured to acquire to-be-recognized images respectively captured by at least two camera devices; the at least two camera devices respectively have different shooting angles; an input module, configured to input the to-be-recognized images into an occlusion recognition model to obtain preferred to-be-recognized images; the preferred to-be-recognized images are images in which the occlusion degree of the driver and passengers is lower than a preset degree; a generation module, configured to input the preferred to-be-recognized images into a behavior recognition model including an attention mechanism layer for processing to obtain the behavior recognition result of the driver and passengers; the attention mechanism layer is configured to increase the weight of the image information in the preferred to-be-recognized images that has an impact on behavior recognition when processing the preferred to-be-recognized images; The driver and passenger behavior recognition apparatus further includes: a splicing module, configured to splice a plurality of the preferred to-be-recognized images to obtain the spliced preferred to-be-recognized images; The splicing module is further configured to: respectively determine the channel dimensions of each of the preferred to-be-recognized images; splice each of the preferred to-be-recognized images in the channel dimension to obtain the spliced preferred to-be-recognized images; the channel dimension of the spliced preferred to-be-recognized images is a preset dimension, and the preset dimension is a multiple of 3.
7. A vehicle, comprising a terminal device, the terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Binocular integrated driver behavior analysis equipment and method thereof
CN107403554A
Personalized fatigue monitoring and reminding system and method
CN109598899A
Recognition method and device and storage medium
CN110633665A
Shielding detection method and device, and electronic equipment
CN113313189A