Training method and device of twin network, computer device and readable storage medium

CN119623549BActive Publication Date: 2026-09-29CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411655870.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2026-09-29
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

[0004]但是,相关的现有技术中,对车辆内部成员的不良行为进行智能检测的准确率都不理想,因此,亟待提供一种能够用于准确识别车辆内部成员行为的行为识别模型,以便于在较短时间内精确识别出车内乘员的不良行为

Benefits of technology

[0040]可见,本申请通过第一样本图像、以及第一样本图像中与第一待识别行为相关的第一目标区域的局部图像(第二样本图像)、以及与第一样本图像中与第二待识别行为相关的第二目标区域的局部图像(第三样本图像),构成样本对集,作为待训练孪生网络的训练样本,使得样本图像同时包括关于样本对象的第一目标区域的图像、第二目标区域的图像,同时还包括样本对象的全局图像,增加了样本图像的类型;其中,全局图像中关联于目标区域(第一目标区域和第二目标区域)的其他特征信息,有利于对目标区域中样本对象的具体行为进行辅助识别,有利于提升后续训练出的孪生网络对于图像中对象的待识别行为的识别准确度;进一步地,本申请将一个样本对集中的样本图像分类为多个类型的样本对,例如第一样本图像和第二样本图像构成第一类型的样本对,第一样本图像和第三样本图像构成第二类型的样本对;基于此,也可将本申请待训练的孪生网络分为包括至少两个子孪生网络,以每一类型的样本对作为对应的一个子孪生网络的输入;并在对孪生网络进行训练的过程中,采用距离相似度和余弦相似度相结合的方式,对一个样本对中的样本图像的第一样本图像特征和第二样本图像特征之间的距离相似度和余弦相似度进行计算,以实现对于该样本对对应的第一损失值的确定,距离相似度和余弦相似度相结合来对第一损失值的确定,有利于提升所得到的第一损失值的准确性,进而有利于提升训练出的孪生网络对于图像中对象的待识别行为的识别精准性;同时,针对于每一样本对,获取其中的第一样本图像特征、第二样本图像特征分别对应的第二损失值、第三损失值;而后,基于第一损失值以及第二损失值、第三损失值,得到各样本对对应的子损失值,进而基于至少第一类型和第二类型的样本对分别对应的子损失值,对样本对集的总损失值进行确定,以基于总损失值的情况实现对于待训练孪生网络的良好训练。综上,基于本申请提供的方法训练出来的孪生网络,能够基于包括车内成员的图片高效且精确地识别出车内乘员关于多个类型的具体行为,例如可以识别出车内成员的一种或两种类型的不良行为,如此有助于及早发现车内成员的不规范或不良的行为,进而提醒车内成员改变相关行为习惯,以有效减少对于驾驶员在行车过程中注意力被分散的问题,从而有利于减少交通事故的发生,以保障驾驶员和乘客的生命安全。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623549B_ABST
    Figure CN119623549B_ABST
Patent Text Reader

Abstract

The application relates to a training method and device of a twin network, a computer device and a readable storage medium. The method comprises the following steps: acquiring a sample pair set; inputting each type of sample pair in the sample pair set into a sub-twin network in a to-be-trained twin network, obtaining first and second sample image features output by a feature extraction network, and obtaining first and second predicted behaviors by a classification network based on the image features; determining a first loss value based on distance similarity and cosine similarity between the first and second sample image features; obtaining second and third loss values based on the first and second predicted behaviors and an actual behavior; determining a sub-loss value of each sample pair based on the first to third loss values, and the sub-loss value of the multi-type sample pair is a total loss value of the sample pair set; in the case that at least one total loss value does not meet a stop training condition, updating network parameters and re-training until a stop condition is met, and obtaining a trained twin network which can be applied to object behavior recognition in a vehicle cabin.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to a method, apparatus, computer device, and computer-readable storage medium for training twin networks. Background Technology

[0002] With the increasing demand for vehicles, traffic accidents occur frequently. Many traffic accidents are related to drivers' poor driving behavior, while others are caused by the poor riding behavior of passengers. In the current vehicle usage environment, drivers often do not pay attention to their driving behavior while driving, and passengers are even less attentive to their riding behavior in the vehicle. For example, it is common for drivers and passengers to smoke or make phone calls in the vehicle. Among these, the driver's phone call while driving is extremely distracting, and the passenger's smoking behavior while driving can easily lead to arguments between the passenger and the driver, which also distracts the driver.

[0003] While vehicles make travel more convenient, the safety of drivers and passengers is paramount. To improve the safety of drivers and passengers, real-time monitoring and intelligent evaluation of driver behavior and passenger behavior can help detect irregular or undesirable behaviors of passengers early on, thereby reminding them to change their habits. This can effectively reduce the problem of drivers being distracted while driving, thus helping to reduce traffic accidents and protect the lives of drivers and passengers.

[0004] However, the accuracy of intelligent detection of undesirable behaviors of vehicle occupants in existing technologies is not ideal. Therefore, there is an urgent need to provide a behavior recognition model that can accurately identify the behavior of vehicle occupants so as to accurately identify undesirable behaviors of vehicle occupants in a short period of time. Summary of the Invention

[0005] Therefore, it is necessary to provide a training method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can train a twin network capable of intelligently monitoring the behavior of occupants inside a vehicle with high accuracy, in order to address the aforementioned technical problems.

[0006] Firstly, this application provides a method for training a Siamese network, the method comprising:

[0007] Obtain a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of a sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified;

[0008] Each type of sample pair is input into the corresponding sub-Twin network in the Siamese network to be trained. Based on the feature extraction network in the sub-Twin network, the first sample image features and the second sample image features corresponding to the two sample images in the sample pair are obtained respectively. The first sample image features and the second sample image features are then input into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior respectively. Different types of sample pairs correspond to different sub-Twin networks.

[0009] A first loss value is determined based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image; a second loss value is obtained based on the first predicted behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second predicted behavior and the actual behavior.

[0010] Based on the first loss value, the second loss value, and the third loss value, a sub-loss value corresponding to each sample pair is determined. Based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs, the total loss value of the sample pair set is determined. If the total loss value does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated, and the Siamese network to be trained is trained again until the stopping training condition is met, resulting in a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0011] In one embodiment, determining the first loss value based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image includes:

[0012] The similarity of sample image features is determined based on the product of the distance similarity and the cosine similarity between the features of the first sample image and the features of the second sample image.

[0013] The sample image feature similarity is converted into a sample image feature similarity probability within a preset value range;

[0014] The first loss value is determined based on the similarity probability of the sample image features, the actual behavior type value of the sample pair, and the preset loss weight.

[0015] In one embodiment, the step of inputting the first sample image features and the second sample image features into the classification network in the sub-Siamese network to obtain the first prediction behavior and the second prediction behavior, respectively, includes:

[0016] The first sample image features are input into the first classification network in the sub-Siamese network to obtain the first prediction behavior;

[0017] The second sample image features are input into the second classification network in the sub-Siamese network to obtain the second prediction behavior.

[0018] In one embodiment, the step of inputting the sample pairs of each type into the corresponding sub-Siamese networks in the Siamese network to be trained, and obtaining the first sample image features and the second sample image features corresponding to the two sample images in the sample pair based on the feature extraction network in the sub-Siamese network, includes:

[0019] For each type of sample pair, one sample image from the sample pair of the type is input into the first feature extraction network in the sub-Siamese network to obtain the corresponding first sample image features;

[0020] Another sample image from the sample pair of the aforementioned type is input into the second feature extraction network in the sub-Siamese network to obtain the corresponding second sample image features.

[0021] In one embodiment, the first feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network;

[0022] The second feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network.

[0023] In one embodiment, each type of sample pair includes both positive and negative sample pairs;

[0024] The two sample images included in the positive sample pair are both images containing the behavior to be identified;

[0025] The negative sample pairs contain two types of sample images that do not contain any behavior to be identified.

[0026] In one embodiment, obtaining the sample pair set includes:

[0027] Acquire multiple sample images including the sample object, and determine the candidate positive sample image and the candidate negative sample image from the multiple sample images;

[0028] For the same behavior to be identified, a first similarity is obtained between the behavior in each candidate positive sample image and the behavior to be identified. Based on the comparison result of the first similarity and the first similarity threshold, positive sample images are selected from multiple candidate positive sample images.

[0029] Obtain a second similarity between the behavior in each of the candidate negative sample images for the same behavior and the behavior to be identified; and filter negative sample images from multiple candidate negative sample images based on the comparison result of the second similarity and the second similarity threshold.

[0030] After image filtering is completed, a second sample image corresponding to the first target region and a third sample image corresponding to the second target region are obtained from each of the filtered positive sample images and negative sample images, respectively; a first type of positive sample pair is formed based on the positive sample image and the corresponding second sample image, a second type of positive sample pair is formed based on the positive sample image and the corresponding third sample image, a first type of negative sample pair is formed based on the negative sample image and the corresponding second sample image, and a second type of negative sample pair is formed based on the negative sample image and the corresponding third sample image;

[0031] The sample pair set is obtained based on the positive sample pairs of the first type, the negative sample pairs of the first type, the positive sample pairs of the second type, and the negative sample pairs of the second type.

[0032] Secondly, this application also provides a training apparatus for Siamese networks, comprising:

[0033] A sample acquisition module is used to acquire a set of sample pairs; the set of sample pairs includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of a sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, wherein the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified;

[0034] A sample processing module is used to input the sample pairs of each type into the corresponding sub-Twin networks in the Siamese network to be trained, obtain first sample image features and second sample image features corresponding to the two types of sample images in the sample pair based on the feature extraction network in the sub-Twin network, and input the first sample image features and second sample image features into the classification network in the sub-Twin network to obtain a first prediction behavior and a second prediction behavior, respectively; wherein, different types of sample pairs correspond to different sub-Twin networks; a first loss value is determined based on the distance similarity and cosine similarity between the first sample image features and the second sample image features; a second loss value is obtained based on the first prediction behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second prediction behavior and the actual behavior.

[0035] The parameter adjustment module is used to determine the sub-loss value corresponding to each sample pair based on the first loss value, the second loss value, and the third loss value; to determine the total loss value of the sample pair set based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs; and to update the network parameters of the Siamese network to be trained and retrain the Siamese network to be trained if the total loss value does not meet the stopping training condition, until the stopping training condition is met, thereby obtaining a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0036] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the steps in the first aspect.

[0037] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the steps in the first aspect.

[0038] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the steps in the first aspect.

[0039] This application provides a training method, apparatus, computer device, computer-readable storage medium, and computer program product for Siamese networks. The training method involves acquiring a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of a sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image; the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified; each type of sample pair is input into a corresponding sub-Siamese network in the Siamese network to be trained; based on the feature extraction network in the sub-Siamese network, the first sample image features and the second sample image features corresponding to the two types of sample images in the sample pair are obtained; and the first sample image features and the second sample image features are then combined. The second sample image features are input into the classification network of the sub-Siamese network to obtain a first predicted behavior and a second predicted behavior, respectively. Different types of sample pairs correspond to different sub-Siamese networks. A first loss value is determined based on the distance similarity and cosine similarity between the first and second sample image features. A second loss value is obtained based on the first predicted behavior and the actual behavior corresponding to the sample pair. A third loss value is obtained based on the second predicted behavior and the actual behavior. Sub-loss values ​​corresponding to each sample pair are determined based on the first, second, and third loss values. The total loss value of the sample pair set is determined based on the sub-loss values ​​corresponding to at least the first and second types of sample pairs. If the total loss value does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated, and the Siamese network to be trained is trained again until the stopping training condition is met, resulting in a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0040] As can be seen, this application constructs a sample pair set using a first sample image, a local image of the first target region in the first sample image related to the first behavior to be identified (a second sample image), and a local image of the second target region in the first sample image related to the second behavior to be identified (a third sample image). This set serves as training samples for the Siamese network to be trained. This ensures that the sample images simultaneously include images of the first target region and the second target region of the sample object, as well as a global image of the sample object, increasing the types of sample images. Furthermore, the other feature information associated with the target regions (first and second target regions) in the global image is beneficial for assisting in the identification of the specific behavior of the sample object in the target region, thus improving the accuracy of the subsequently trained Siamese network in identifying the behavior of the object in the image. Further, this application classifies the sample images in a sample pair set into multiple types of sample pairs. For example, the first sample image and the second sample image constitute a first type of sample pair, and the first sample image and the third sample image constitute a second type of sample pair. Based on this, the Siamese network to be trained in this application can also be divided into... The method includes at least two sub-Siamese networks, with each type of sample pair serving as the input to a corresponding sub-Siamese network. During the training of the Siamese network, a combination of distance similarity and cosine similarity is used to calculate the distance similarity and cosine similarity between the first and second sample image features of a sample image in a sample pair. This determines the first loss value for that sample pair. The combination of distance similarity and cosine similarity in determining the first loss value improves the accuracy of the obtained first loss value, thereby enhancing the accuracy of the trained Siamese network in recognizing the behavior of objects in the image. Simultaneously, for each sample pair, the second and third loss values ​​corresponding to the first and second sample image features are obtained, respectively. Then, based on the first, second, and third loss values, the sub-loss values ​​corresponding to each sample pair are obtained. Finally, based on the sub-loss values ​​corresponding to at least the first and second types of sample pairs, the total loss value of the sample pair set is determined, thus achieving good training of the Siamese network to be trained based on the total loss value. In summary, the twin network trained based on the method provided in this application can efficiently and accurately identify multiple types of specific behaviors of vehicle occupants based on images including those of vehicle occupants. For example, it can identify one or two types of undesirable behaviors of vehicle occupants. This helps to detect irregular or undesirable behaviors of vehicle occupants early, thereby reminding them to change their related behavioral habits. This effectively reduces the problem of driver distraction during driving, thus helping to reduce the occurrence of traffic accidents and protect the lives of drivers and passengers. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a diagram illustrating the application environment for training a twin network in one embodiment.

[0043] Figure 2 This is a diagram illustrating the application environment of a trained twin network in one embodiment.

[0044] Figure 3 This is a flowchart illustrating a training method for a Siamese network in one embodiment;

[0045] Figure 4 This is a schematic diagram of the structure of a twin network in one embodiment;

[0046] Figure 5 This is a schematic diagram of the twin network structure in another embodiment;

[0047] Figure 6 This is a structural block diagram of a training device for a twin network in one embodiment;

[0048] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0050] The training method for Siamese networks provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 902 communicates with server 904 via a network. Terminal 902 can at least acquire sample images for training the Siamese network to be trained, and perform certain preprocessing operations on the sample images to obtain the first sample image corresponding to the sample image; in addition, terminal 902 can also perform classification labeling of the sample images. The data storage system can store the data that server 904 needs to process. The data storage system can be integrated on server 904, or it can be placed on the cloud or other network servers. The data storage system can at least store multiple sample pairs including the first sample image, the second sample image, and the third sample image, as well as the Siamese network to be trained and the trained Siamese network. Terminal 902 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. Server 904 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0051] It should be noted that the training process of the twin network can be based on the server 904 providing the computation, or it can be based on the terminal 902 providing the computation, or it can be based on the communication between the terminal 902 and the server 904 to achieve the computation. This application does not make any specific limitation in this regard.

[0052] Furthermore, the trained Siamese network provided in this application embodiment can be applied to, for example... Figure 2 In the application environment shown. But Figure 2 This application provides only one possible application environment for the trained Siamese network, and is not limited thereto; that is, the trained Siamese network can be used not only for applications such as... Figure 2 The driver / occupant behavior recognition shown in the vehicle environment can also be used for, for example, worker behavior recognition in a factory environment, and student behavior recognition in a teaching environment, etc. Specifically, this application scenario includes a vehicle 100, which includes a vehicle controller 101, and an image acquisition device 102 is included in the vehicle cabin.

[0053] The vehicle controller 101 can first acquire images of target objects (such as at least one of the driver and passengers) in the cabin through the image acquisition device 102, and then process the acquired images to obtain a first image of the target object and store the first image. In the process of recognizing object behavior in an image using the Siamese network trained according to this application, the vehicle controller 101 can first acquire a first image of the target object, and determine a second image and a third image based on the first image. That is, the first local image (second image) and the second local image (third image) of the target object are obtained from the global image (first image) of the target object. The first local image is the image corresponding to the region in the first image that is related to the first behavior to be recognized of the target object, and the second local image is the image corresponding to the region in the first image that is related to the second behavior to be recognized of the target object. Further, the vehicle controller 101 inputs the first image, the second image and the third image belonging to the same sample pair set into the Siamese network trained according to this application to obtain the image feature similarity probability between the two images in each sample pair based on the Siamese network trained according to this application. Finally, based on the image feature similarity probabilities corresponding to the sample pairs of multiple types, the recognition results for the first behavior to be recognized and the second behavior to be recognized of the target object are determined to determine the specific behavior of the target object.

[0054] In order to provide a behavior recognition model that can accurately identify the behavior of occupants inside a vehicle, so as to accurately identify the bad behavior of occupants in a short time, this application provides a method for training a Siamese network to train a Siamese network that can identify the behavior of occupants inside a vehicle with high accuracy.

[0055] In one exemplary embodiment, such as Figure 3 As shown, a method for training a Siamese network is provided, including but not limited to the following steps S201 to S204. Wherein:

[0056] S201, Obtain a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of the sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified.

[0057] The sample objects can be either the driver or the occupants inside the vehicle. The first sample image can be an image containing at least one of the driver and occupants. If the first behavior to be identified is whether the driver is smoking, then the first target region is the region related to smoking, such as the driver's mouth area. In this case, the second sample image is a local image of the mouth area in the driver's first sample image. If the second behavior to be identified is whether the driver is making a phone call, then the second target region is the region related to making a phone call, such as the driver's eyes area. In this case, the third sample image is a local image of the eyes area in the driver's first sample image. The local image corresponding to the target region is also known as the face ROI image. ROI (Region of Interest) refers to a specific area in an image that is being processed, usually delineated using rectangles, circles, ellipses, irregular polygons, etc. That is, the second and third sample images can be, for example, images of two different regions bounded by rectangles in the first sample image.

[0058] The first sample image obtained in this application and the second sample image associated with the first sample image can form a first type of sample pair, such as a sample involving smoking behavior. The first sample image obtained in this application and the third sample image associated with the first sample image can form a second type of sample pair, such as a sample involving telephone behavior. A sample pair set can be constructed using the first type of sample pairs and the second type of sample pairs. It should be noted that the embodiment provided herein, which includes both first and second type sample pairs in the sample pair set, is merely one optional implementation method provided by this application. However, this application is not limited to this, and a sample pair set may also include three or more types of sample pairs simultaneously.

[0059] In one optional implementation, before step S201, an image including the sample object can be acquired by an image acquisition device inside the vehicle cabin, and the image can be preprocessed to obtain the processed image including the sample object as the first sample image. The image preprocessing may include, for example, denoising or cropping, and this application does not specifically limit the specific methods used.

[0060] The image acquisition device can be, for example, an infrared sensor or a visible light sensor. An infrared sensor, such as an infrared camera, is used to acquire infrared images, while a visible light sensor, such as a regular camera, is used to acquire color images, also known as RGB images, where R represents the red channel, G the green channel, and B the blue channel. The image of the sample object can be a single-channel infrared image of the vehicle cabin including the sample object, acquired by an infrared sensor, or an RGB color image of the vehicle cabin including the sample object, acquired by a visible light sensor. The image in the infrared image sample pair can be an image formed based on acquiring the radiation of the sample object in the infrared band.

[0061] This application provides an optional implementation in which the vehicle cabin includes both infrared sensors and visible light sensors, wherein the infrared sensors are used to monitor and photograph the driver's seat in the vehicle, and the visible light sensors are used to monitor and photograph the passenger seats in the vehicle; however, this is only one optional implementation provided by this application, and this application is not limited thereto.

[0062] Furthermore, an alternative implementation is provided in which, if the obtained image of the sample object is a single-channel infrared image, it can be further converted into a three-channel infrared image, so as to use the three-channel infrared image as the first sample image of the sample object.

[0063] Each sample pair set includes three images: a first sample image containing the sample object, a second sample image cropped from the first sample image based on a region of the sample object related to a first behavior to be identified, and a third sample image cropped from the first sample image based on a region of the sample object related to a second behavior to be identified. Further, the second sample image and its corresponding first sample image constitute one type of sample pair, and the third sample image and its corresponding first sample image constitute another type of sample pair.

[0064] S202, each type of sample pair is input into the corresponding sub-Twin network in the Siamese network to be trained, and the first sample image feature and the second sample image feature corresponding to the two sample images in the sample pair are obtained based on the feature extraction network in the sub-Twin network; the first sample image feature and the second sample image feature are input into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior respectively; wherein, different types of sample pairs correspond to different sub-Twin networks.

[0065] Specifically, the Siamese network (Siamese neural network) to be trained provided in this application may include a feature extraction network and a classification network. That is, the Siamese network is composed of at least a feature extraction network and a classification network. The input data of the feature extraction network is a set of sample pairs, and the input data of the classification network is the data output by the feature extraction network after processing the set of sample pairs.

[0066] In step S202 of this application, the second sample image and its corresponding first sample image are combined to form a sample pair of, for example, a first type, and the third sample image and its corresponding first sample image are combined to form a sample pair of, for example, a second type. Since a sample pair is simultaneously input into the Siamese network to be trained as a set of data, this application provides an alternative implementation method in which, based on the sample pair set including first type and second type sample pairs, the Siamese network to be trained is divided into two sub-Siamese networks, each of which is provided with a corresponding feature extraction network and a classification network. For example, it includes a first sub-Siamese network for receiving first type sample pairs and a second sub-Siamese network for receiving second type sample pairs.

[0067] In one optional implementation, a sample pair of a certain type, comprising two sample images (e.g., a second sample image and a first sample image, or a third sample image and a first sample image), is input into a sub-Siamese network of the Siamese network to be trained provided in this application. The feature extraction network in the sub-Siamese network receives the sample pair to obtain first and second sample image features output by the feature extraction network after processing the two sample images respectively. Then, the first and second sample image features are input into a classification network included in the sub-Siamese network to obtain a first prediction based on the classification network's output after processing the first sample image features, and a second prediction based on the classification network's output after processing the second sample image features.

[0068] In an exemplary embodiment, the first sample image feature can be an image feature vector, and the second sample image feature can also be an image feature vector; for example, features in the first sample image (e.g., the whole image) are extracted by a feature extraction network to obtain a first feature vector (first sample image feature), and features in the second sample image (e.g., the ROI image) are extracted by a feature extraction network to obtain a second feature vector (second sample image feature).

[0069] In an exemplary embodiment, the classification network is used to receive first sample image features corresponding to one sample image in a sample pair, and outputs a prediction result (first predicted behavior) of the specific behavior of the sample object in the sample image; the classification network is also used to receive second sample image features corresponding to the other sample image in the sample pair, and outputs a prediction result (second predicted behavior) of the specific behavior of the sample object in the sample image. For example, if the first sample image includes smoking behavior, then the corresponding first predicted behavior may be smoking behavior; if the sample image includes phone call behavior, then the corresponding first predicted behavior may be phone call behavior; the second predicted behavior is similar and will not be described in detail here.

[0070] It should also be noted that the classification network makes predictions based on pre-defined types of behaviors to be identified. For example, if the pre-defined types of behaviors to be identified include smoking, the classification network will output that the first predicted behavior is smoking.

[0071] S203, determine a first loss value based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image; obtain a second loss value based on the first predicted behavior and the actual behavior corresponding to the sample pair; and obtain a third loss value based on the second predicted behavior and the actual behavior.

[0072] For calculating distance similarity, Euclidean distance can be used to calculate the similarity of images. The smaller the Euclidean distance, the greater the similarity. For calculating cosine similarity, the cosine value of the angle between two feature vectors is used to measure the cosine similarity between the two feature vectors. The more similar the two vectors are, the smaller the angle between them, and the closer the cosine value is to 1.

[0073] The first predicted behavior obtained in step S202 of this application is a prediction result of the behavior to be identified of a sample object in a sample image of a sample pair, based on the feature extraction network and classification network in the sub-Siamese network to be trained provided in this application. The user can also determine the actual behavior of the sample object included in the sample image through expert recognition (human recognition). Then, based on the difference between the first predicted behavior for the same sample image and the actual behavior corresponding to the sample pair, for example, by data processing of the numerical value corresponding to the behavior type of the first predicted behavior and the numerical value corresponding to the behavior type of the actual behavior, a second loss value for the first predicted behavior and the actual behavior can be obtained. The second predicted behavior obtained in step S202 of this application is a prediction result of the behavior to be identified in the sample object in another sample image of a sample pair, based on the feature extraction network and classification network in the sub-Siamese network of the Siamese network to be trained provided in this application. The user can also determine the actual behavior of the sample object included in the sample image, and then, based on the difference between the second predicted behavior for the same sample image and the corresponding actual behavior of the sample pair, for example, through data processing of the numerical values ​​corresponding to the behavior type of the second predicted behavior and the numerical values ​​corresponding to the behavior type of the actual behavior, a third loss value for the second predicted behavior and the actual behavior can be obtained. Simultaneously with obtaining the second and third loss values, a first loss value corresponding to the first sample image features and the second sample image features can also be obtained based on the distance similarity and cosine similarity between the first sample image features and the second sample image features. The calculation of the loss value can be based on a loss function.

[0074] S204. Based on the first loss value, the second loss value, and the third loss value, determine the sub-loss value corresponding to each sample pair. Based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs respectively, determine the total loss value of the sample pair set. If the total loss value does not meet the stopping training condition, update the network parameters of the Siamese network to be trained, and retrain the Siamese network to be trained until the stopping training condition is met, to obtain the trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0075] In this process, after obtaining the first, second, and third loss values ​​for a sample pair of a certain type based on execution step S203, sub-loss values ​​for that sample pair can be further determined based on these three loss values. Then, in the same way, sub-loss values ​​corresponding to sample pairs of multiple types included in the sample pair set can be obtained respectively; for example, the sub-loss values ​​of the first type of sample pair and the second type of sample pair are obtained. Then, the total loss value corresponding to the sample pair set can be determined based on these two sub-loss values. Then, it is determined whether the total loss value of the sample pair set meets the stopping training condition of the Siamese network to be trained. If the total loss value of some samples does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated based on the total loss value of each sample pair set used in a training process, and the Siamese network is trained again until the total loss values ​​of multiple sample pair sets all meet the stopping training condition. At this time, the training process of the Siamese network can be ended, and the trained Siamese network can be obtained. Since the data input to the Siamese network to be trained in this application is a vehicle cabin image including sample objects, and at least some of the sample objects included in the vehicle cabin image have behaviors to be identified, the Siamese network trained in this application can be used to identify the (pre-set) behaviors of objects in the vehicle cabin; thus, when identifying the behavior of objects in the vehicle cabin based on the trained Siamese network in the future, the accuracy of the identification of the behavior of objects in the vehicle cabin can be improved.

[0076] It should be noted that the sample set provided above is an embodiment including cabin images of sample objects, and is only one optional implementation method provided by this application. This application is not limited thereto. The specific image type including sample objects used to train the Siamese network can be selected based on the different application scenarios required by the trained Siamese network. For example, the sample set can be selected as classroom images including sample objects (a student). In this case, the corresponding preset behavior can be sleeping behavior, face-covering behavior, eating behavior, etc.

[0077] The above-mentioned training method for Siamese networks involves acquiring a set of sample pairs. This set includes at least two types of sample pairs: a first type and a second type. The first type of sample pair includes a first sample image and a second sample image of the sample object; the second type of sample pair includes a first sample image and a third sample image. The second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image. The first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified. Each type of sample pair is input into its corresponding sub-Siamese network within the Siamese network to be trained. Based on the feature extraction network in the sub-Siamese network, the first sample image features and the second sample image features corresponding to the two types of sample images in the sample pair are obtained. These first and second sample image features are then input into the classification network in the sub-Siamese network for further processing. The process involves obtaining a first predicted behavior and a second predicted behavior, where different types of sample pairs correspond to different sub-Siamese networks; determining a first loss value based on the distance similarity and cosine similarity between the features of the first and second sample images; obtaining a second loss value based on the first predicted behavior and the actual behavior corresponding to the sample pair; and obtaining a third loss value based on the second predicted behavior and the actual behavior; determining sub-loss values ​​corresponding to each sample pair based on the first, second, and third loss values; determining the total loss value of the sample pair set based on the sub-loss values ​​corresponding to at least the first and second types of sample pairs; and updating the network parameters of the Siamese network to be trained if the total loss value does not meet the stopping training condition, and retraining the Siamese network to be trained until the stopping training condition is met, thereby obtaining a trained Siamese network. The trained Siamese network is then applied to the recognition of at least two behaviors of objects within the vehicle cabin.

[0078] As can be seen, this application constructs a sample pair set using a first sample image, a local image of the first target region in the first sample image related to the first behavior to be identified (a second sample image), and a local image of the second target region in the first sample image related to the second behavior to be identified (a third sample image). This set serves as training samples for the Siamese network to be trained. This ensures that the sample images simultaneously include images of the first target region and the second target region of the sample object, as well as a global image of the sample object, increasing the types of sample images. Furthermore, the other feature information associated with the target regions (first and second target regions) in the global image is beneficial for assisting in the identification of the specific behavior of the sample object in the target region, thus improving the accuracy of the subsequently trained Siamese network in identifying the behavior of the object in the image. Further, this application classifies the sample images in a sample pair set into multiple types of sample pairs. For example, the first sample image and the second sample image constitute a first type of sample pair, and the first sample image and the third sample image constitute a second type of sample pair. Based on this, the Siamese network to be trained in this application can also be divided into... The method includes at least two sub-Siamese networks, with each type of sample pair serving as the input to a corresponding sub-Siamese network. During the training of the Siamese network, a combination of distance similarity and cosine similarity is used to calculate the distance similarity and cosine similarity between the first and second sample image features of a sample image in a sample pair. This determines the first loss value for that sample pair. The combination of distance similarity and cosine similarity in determining the first loss value improves the accuracy of the obtained first loss value, thereby enhancing the accuracy of the trained Siamese network in recognizing the behavior of objects in the image. Simultaneously, for each sample pair, the second and third loss values ​​corresponding to the first and second sample image features are obtained, respectively. Then, based on the first, second, and third loss values, the sub-loss values ​​corresponding to each sample pair are obtained. Finally, based on the sub-loss values ​​corresponding to at least the first and second types of sample pairs, the total loss value of the sample pair set is determined, thus achieving good training of the Siamese network to be trained based on the total loss value. In summary, the twin network trained based on the method provided in this application can efficiently and accurately identify multiple types of specific behaviors of vehicle occupants based on images including those of vehicle occupants. For example, it can identify one or two types of undesirable behaviors of vehicle occupants. This helps to detect irregular or undesirable behaviors of vehicle occupants early, thereby reminding them to change their related behavioral habits. This effectively reduces the problem of driver distraction during driving, thus helping to reduce the occurrence of traffic accidents and protect the lives of drivers and passengers.

[0079] In an exemplary embodiment, the step S203, which determines the first loss value based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image, can specifically be executed as follows: steps S231-S233 (not shown), wherein:

[0080] S231, determine the similarity of sample image features based on the product of the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image;

[0081] S232, convert the sample image feature similarity into the sample image feature similarity probability within a preset value range;

[0082] S233, determine the first loss value based on the similarity probability of sample image features, the actual behavior type value of the sample pair, and the preset loss weight.

[0083] Specifically, firstly, based on the formula for calculating distance similarity, such as Euclidean distance, the distance similarity between the features of the first sample image and the features of the second sample image is calculated; then, based on the formula for calculating cosine similarity, the cosine similarity between the features of the first sample image and the features of the second sample image is calculated; and finally, based on the product of the distance similarity and the cosine similarity corresponding to the same sample pair, the similarity of the sample image features corresponding to the two images in the sample pair is calculated.

[0084] Then, the similarity of the sample image features is converted into the similarity probability of the sample image features within a preset value range. For example, the similarity of the sample image features is converted into the relative probability of the sample image features within the range of [0,1]. The relative probability of the sample image features is used to determine whether the behavior of the two sample images in the sample pair is the same. For example, if the relative probability of the sample image features is greater than the preset smoking probability, then the judgment result is that the sample objects in the two images of the sample pair both have smoking behavior.

[0085] Finally, based on the sample image feature similarity probability obtained in step S232, the actual behavior type value of the sample objects in the sample pair image, and the preset loss weight, the first loss value associated with a sample pair can be determined. The actual behavior type value, for example, can be set to 1 if the actual behavior is the same as the preset detection behavior, and set to 0 if the actual behavior is different from the preset detection behavior.

[0086] In an exemplary embodiment, the step S202 described above, which involves inputting the first sample image features and the second sample image features into the classification network of the sub-Siamese network to obtain the first prediction behavior and the second prediction behavior, can be optionally executed as follows: inputting the first sample image features into the first classification network of the sub-Siamese network to obtain the first prediction behavior; and inputting the second sample image features into the second classification network of the sub-Siamese network to obtain the second prediction behavior.

[0087] In an exemplary embodiment, the step S202 described above, which involves inputting sample pairs of each type into the corresponding sub-Siamese networks in the Siamese network to be trained, and obtaining the first sample image features and the second sample image features corresponding to the two sample images in the sample pair based on the feature extraction network in the sub-Siamese network, can be optionally executed as follows: for each type of sample pair, inputting one sample image from the sample pair of the type into the first feature extraction network in the sub-Siamese network to obtain the corresponding first sample image features; and inputting the other sample image from the sample pair of the type into the second feature extraction network in the sub-Siamese network to obtain the corresponding second sample image features.

[0088] Specifically, one alternative implementation is provided in which the feature extraction network included in the sub-Siamese network to be trained provided in this application specifically includes a first feature extraction network and a second feature extraction network, and the classification network included in the Siamese network to be trained specifically includes a first classification network and a second classification network; wherein, the first classification network is used to receive the output result of the corresponding first feature extraction network, and the second classification network is used to receive the output result of the corresponding second feature extraction network.

[0089] That is, the Siamese network to be trained provided in this application can be considered to include two sub-Siamese networks. The first sub-Siamese network includes a first feature extraction network and a first classification network, as well as a second feature extraction network and a second classification network. The first sub-Siamese network is used to receive and process the two images in a sample pair.

[0090] For example, the first sample image in a sample pair is input to a first feature extraction network in a sub-Siamese network. The first feature extraction network outputs features of the first sample image, which are then input to a first classification network. Based on the corresponding first classification network, a first predicted behavior for the sample object in the first sample image is obtained. The second sample image in the sample pair is input to a second feature extraction network in the sub-Siamese network. The second feature extraction network outputs features of the second sample image, which are then input to a second classification network. Based on the corresponding second classification network, a second predicted behavior for the sample object in the second sample image is obtained. For example, if the sample image in the sample pair contains smoking behavior, the sample pair is input to a sub-Siamese network of the Siamese network to be trained provided in this application. The first and second prediction results output by the classification network therein may both be "smoking behavior".

[0091] It should also be added that, regarding the calculation of the second and third loss values, appropriate loss functions for image classification can be selected for the relevant calculations. For example, the first loss function used to calculate the second loss value and the second loss function used to calculate the third loss value can include, but are not limited to, the mean squared error loss function, the mean absolute error loss function, or the cross-entropy loss function. The first and second loss functions can be the same or different. For example, the loss function used to calculate both the second and third loss values ​​can be the cross-entropy loss function, which calculates the loss value between the predicted behavior type value and the actual behavior type value for the corresponding set of predicted behaviors.

[0092] In an exemplary embodiment, the first feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network; the second feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network.

[0093] Specifically, each sub-Siamese network of the Siamese network to be trained provided in this application includes a first feature extraction network and a second feature extraction network. Both of these networks can be selected to use convolutional neural networks. The specific type of convolutional neural network can be, for example, either the lightweight neural network MobileNet or the object detection network YOLO.

[0094] It should also be noted that for the first and second feature extraction networks included in each sub-Siamese network, the same type of convolutional neural network can be chosen, or different types of convolutional neural networks can be chosen. For example, when the first feature extraction network is MobileNet-V3, the second feature extraction network is also MobileNet-V3; when the first feature extraction network is MobileNet-V3, the second feature extraction network is YOLO-V8; when the first feature extraction network is YOLO-V8, the second feature extraction network is MobileNet-V3; and when the first feature extraction network is YOLO-V8, the second feature extraction network is also YOLO-V8.

[0095] The MobileNet-V3 and YOLO-V8 provided here are optional versions of the MobileNet and YOLO convolutional neural networks, respectively. However, this application does not limit the specific version selection and can be adjusted and used according to needs.

[0096] Specifically, in the Siamese network to be trained provided in this application, if it includes two sub-Siamese networks, the specific types of the two first feature extraction networks and the two second feature extraction networks included in the two sub-Siamese networks can be selected from MobileNet-V3 and YOLO-V8, respectively.

[0097] In an exemplary embodiment, each type of sample pair includes a positive sample pair and a negative sample pair; the two sample images contained in the positive sample pair are images with the behavior to be identified; the two sample images contained in the negative sample pair are images without the behavior to be identified.

[0098] Specifically, the sample pair set collected in this application for training the Siamese network includes both positive and negative sample pairs. In positive sample pairs, the behavior of the sample objects in the images includes a pre-defined behavior to be identified. For example, if the pre-defined behavior is smoking, then the sample objects (e.g., drivers) in the first and second sample images of the same sample pair will both exhibit smoking behavior. In negative sample pairs, the behavior of the sample objects does not include the pre-defined behavior to be identified. For example, if the pre-defined behavior is smoking, then the sample objects (e.g., drivers) in the first and second sample images of the same sample pair will not exhibit smoking behavior.

[0099] By selecting training sample pairs for training the Siamese network, including both positive and negative sample pairs, it is beneficial to improve the recognition accuracy of the finally trained Siamese network for the behavior to be identified.

[0100] It should also be noted that applying the trained Siamese network to object behavior recognition within a vehicle cabin is only one possible application scenario provided by this application. However, in other similar scenarios, the Siamese network trained using the training method provided in this application can also be used to recognize the behavior of objects within that scenario. For example, the Siamese network trained using the training method provided in this application can be used to recognize worker behavior in a factory environment; or it can be used to recognize student behavior in a classroom environment. Depending on the application scenario of the Siamese network, the collection scenario of its training samples can also be adapted to the corresponding scenario, which is beneficial to improving the behavior recognition performance of the trained Siamese network.

[0101] In an exemplary embodiment, obtaining the sample pair set in step S201 can be specifically achieved by executing steps S211-S215 (not shown), wherein:

[0102] S211, acquire multiple sample images including the sample object, and determine the candidate positive sample image and the candidate negative sample image from the multiple sample images;

[0103] S212, for the same behavior to be identified, obtain the first similarity between the behavior in each candidate positive sample image and the behavior to be identified, and filter positive sample images from multiple candidate positive sample images based on the comparison result of the first similarity and the first similarity threshold.

[0104] S213, obtain the second similarity between the behavior and the behavior to be identified in each candidate negative sample image for the same behavior, and filter negative sample images from multiple candidate negative sample images based on the comparison result of the second similarity and the second similarity threshold.

[0105] S214, after completing the image filtering, obtain the second sample image corresponding to the first target region and the third sample image corresponding to the second target region from each of the filtered positive sample images and negative sample images; form a first type of positive sample pair based on the positive sample image and the corresponding second sample image, form a second type of positive sample pair based on the positive sample image and the corresponding third sample image, form a first type of negative sample pair based on the negative sample image and the corresponding second sample image, and form a second type of negative sample pair based on the negative sample image and the corresponding third sample image;

[0106] S215, based on the first type of positive sample pairs, the first type of negative sample pairs, the second type of positive sample pairs, and the second type of negative sample pairs, a sample pair set is obtained.

[0107] Specifically, firstly, multiple sample images containing sample objects are acquired, and then these sample images are divided into candidate positive sample images and candidate negative sample images. Next, the first similarity between the behavior of the sample object in each candidate positive sample image and the behavior to be identified is calculated, i.e., the similarity probability between the behavior of the sample object in the candidate positive sample image and the behavior to be identified. Then, based on the comparison result of the first similarity and a first similarity threshold, candidate positive sample images with a first similarity greater than or equal to the first similarity threshold are selected as the selected positive sample images. Simultaneously, the second similarity between the behavior of the sample object in each candidate negative sample image and the behavior to be identified is calculated, i.e., the similarity probability between the behavior of the sample object in the candidate negative sample image and the behavior to be identified. Then, based on the comparison result of the second similarity and a second similarity threshold, candidate negative sample images with a second similarity less than the second similarity threshold are selected as the selected negative sample images.

[0108] It should be added that this application does not limit the type of specific calculation formula used for the behavior of sample objects in sample images and the similarity probability between them and the behavior to be identified. Appropriate similarity calculation formulas can be selected according to the needs.

[0109] After image filtering is completed, corresponding second and third sample images are extracted from each positive and negative sample image. Then, a first type of positive sample pair is formed based on the positive sample image and the corresponding second sample image, and a second type of positive sample pair is formed based on the positive sample image and the corresponding third sample image. Similarly, a first type of negative sample pair is formed based on the negative sample image and the corresponding second sample image, and a second type of negative sample pair is formed based on the negative sample image and the corresponding third sample image. This process creates a sample pair set by including multiple types of positive and negative sample pairs.

[0110] It should also be added that when the selected sample set includes vehicle cabin images of the sample objects, the images used to train the Siamese network to be trained must include vehicle cabin images, but it is not limited to all images being vehicle cabin images; other images may also be included.

[0111] Furthermore, it should be added that, for a Siamese network applied to the same application scenario, the sample images of the sample pair set used when training the Siamese network can be collected in the same application scenario or in multiple application scenarios. This application does not make any specific limitations on this.

[0112] Regarding the acquisition of the sample pair set in step S201, this application also provides an alternative implementation method as follows:

[0113] A first sample image of a sample object inside the vehicle cabin is acquired; the first sample image is input into a face detection model, and if the first sample image contains a sample object, the face information corresponding to the sample object is output; based on the face information, the face image corresponding to the sample object is extracted from the first sample image; the face image is input into a face key point detection model to obtain the face key point coordinates of the sample object; based on the face key point coordinates, a first target region and a second target region are determined, and a second sample image corresponding to the first target region and a third sample image corresponding to the second target region are obtained from the face image; the first sample image, the second sample image, and the third sample image of the sample object constitute a sample pair set.

[0114] Regarding the training method for Siamese networks applied to object behavior recognition in vehicle cockpits provided in this application, this application provides an optional embodiment, and the specific training process is shown in steps S1-S7 below, wherein:

[0115] S1, obtain multiple sample pairs; each sample pair includes a full-image smoking or phone call sample image (first sample image), a mouth ROI smoking sample image (second sample image), and an eye ROI phone call sample image (third sample image), wherein the full-image smoking or phone call sample image is in Figure 5 In the image, D1 represents the mouth ROI (Region of Interest) in a smoking sample image. Figure 5 In the image, D2 represents the smoking sample image of the eye ROI. Figure 5 The image is denoted by D3. For the same sample pair, the smoking sample image D2 (mouth ROI) is obtained by cropping the region related to the smoking behavior to be identified from the full-image smoking or phone call sample image D1, and the phone call sample image D3 (eye ROI) is obtained by cropping the region related to the phone call behavior to be identified from the full-image smoking or phone call sample image D1.

[0116] S2, using convolutional neural network 2 to extract the ROI features of dataset D2, yielding the feature vector G. w (D2), which is the second sample image feature mentioned above; the ROI features of dataset D3 are extracted using convolutional neural network 4 to obtain the feature vector G. w (D3), which is the second sample image feature mentioned above; the full image features of dataset D1 are extracted using convolutional neural network 1, and the full image features of dataset D1 are extracted using convolutional neural network 3, resulting in feature vector G. w(D1), which refers to the first sample image features mentioned above; among them, the four convolutional neural networks 1 to 4 can all adopt either the lightweight neural network MobileNet-V3 or the object detection network YOLO-V8, depending on the requirements. It should be noted that if convolutional neural network 1 and convolutional neural network 3 are of different types, for example, if convolutional neural network 1 uses MobileNet-V3 and convolutional neural network 3 uses YOLO-V8, even if convolutional neural network 1 and convolutional neural network 3 receive the same dataset D1, the first sample image features obtained after processing by convolutional neural network 1 and convolutional neural network 3 respectively may differ. It should also be noted that, at least for samples of the same type, the corresponding networks can share parameters.

[0117] The subsequent steps S3-S7, to Figure 5 The convolutional neural network 3 shown in the image has D4 as input on the left and G as output on the right. w (D4) will be used as an example for explanation.

[0118] S3 calculates the feature vector G using a combination of distance similarity and cosine similarity metrics. w (D1) and eigenvector G w The similarity measure of (D2) is the feature similarity of the first sample image. The specific calculation formula is shown in formula (1).

[0119] S(G w (D1), G w (D2))=dis(G w (D1), G w (D2))·cos(G w (D1), G w (D2)) (1)

[0120] In formula (1), dis(G) w (D1),G w (D2)) represents G w (D1) and G w The distance similarity between (D2) and cos(G) w (D1),G w (D2)) represents G w (D1) and G w Cosine similarity between (D2); where G w (D1) is, for example, the first feature vector, representing the features of the first sample image; G w(D2) For example, the second feature vector represents the features of the second sample image; dis() represents the distance similarity; cos() represents the cosine similarity.

[0121] The feature vector G is calculated by combining distance similarity and cosine similarity. w (D3) and eigenvector G w The similarity measure of (D4) is also known as the feature similarity of the second sample image. The specific calculation formula is shown in formula (2).

[0122] S(G w (D3), G w (D4))=dis(G w (D3), G w (D4))·cos(G w (D3), G w (D4)) (2)

[0123] In formula (2), dis(G) w (D3),G w (D4) represents G w (D3) and G w The distance similarity between (D4) and cos(G) w (D3),G w (D4) represents G w (D3) and G w Cosine similarity between (D4); where G w (D4) For example, the first feature vector represents the features of the first sample image; G w (D3) is, for example, the second feature vector, which represents the features of the second sample image.

[0124] S4, the feature vector G w (D1), Eigenvector G w (D2) Similarity is processed through a fully connected layer and then a Sigmoid layer, transforming it into the relative probability (RPS) of numerical samples within the [0,1] interval, i.e., the similarity probability of judging a sample as smoking. When RPS > 0.5, it indicates that the model judges all samples as smoking, showing that the ROI image classification model (second classification network) is more consistent in its judgment of positive smoking samples in the mouth ROI with the global image classification task (first classification network) in its judgment of positive smoking samples in the whole image. In particular, during the training of the Siamese network (model), the model tends to judge positive smoking samples.

[0125] Eigenvector G w (D3), Eigenvector G w(D4) The similarity score is processed through a fully connected layer and then a Sigmoid layer, transforming it into the relative probability (RPS) of numerical samples within the [0,1] interval. This RPS represents the probability that a sample is a phone call. When RPS > 0.5, it indicates that the model classifies the sample as a phone call, demonstrating a high degree of consistency between the ROI image classification model (second classification network) and the global image classification task (first classification network) in classifying positive phone call samples across the entire image. Specifically, during Siamese network (model) training, the model tends to classify these samples as positive phone call samples.

[0126] In other words, the higher the probability of similarity of sample image features, the more consistent the recognition result of the first classification network in the Siamese network for the target object's behavior in the first sample image (full image) is with the recognition result of the second classification network for the target object's behavior in the second sample image (partial image).

[0127] S5. Construct the cross-entropy loss function Lrps based on RPS. The relevant construction formula is shown in formula (3).

[0128]

[0129] Wherein, L in formula (3) rps λ represents the first loss value, which is the loss value between the first sample image and the second sample image; RPS represents the sample image feature similarity probability, which is the relative probability of the sample image feature vector being converted into a numerical sample within a preset interval; λ represents the loss weight (preset loss weight); l represents the label of the sample group, which is the actual behavior type of the sample group. If the sample group is a positive sample group, l is 1, and if the sample group is a negative sample group, l is 0. This represents a regularization function, which can enhance the generalization ability of the model.

[0130] In training the model (Siamese network), similarity measurement and classification tasks can learn each other, enabling the model to distinguish whether two samples are similar and to identify the category of a single sample. In addition, since ROI image classification tasks and full-image image classification tasks have different noise patterns, learning the two image classification tasks simultaneously can average the noise, share the risk of overfitting, and obtain a more generalized representation.

[0131] S6, the sub-Siamese network of the model includes two classification tasks (one is the ROI image classification task, and the other is the whole image classification task) and a similarity measurement task. Therefore, the sub-loss function consists of three parts, that is, the sub-loss value consists of the first loss value, the second loss value, and the third loss value. The calculation method of the sub-loss value is shown in formulas (4) and (5).

[0132] L1=λ×(L cls (G w (D1)+L cls (G w (D2)+(1-2λ)×L rps (4)

[0133] L2=λ×(L cls (G w (D3)+L cls (G w (D4)+(1-2λ)×L rps (5)

[0134] In formulas (4) and (5), L1 and L2 represent the sub-loss values ​​of the corresponding sample pairs; λ represents the loss weight; G w (D1), G w (D4) are all the first feature vectors, representing the features of the first sample image; G w (D2), G w (D3) is the second feature vector, representing the features of the second sample image; L rps This represents the first loss value, which is the loss value between two images in a sample pair of the same type.

[0135] S7. Based on the sum of the sub-loss values ​​L1 corresponding to the first type of sample pair and L2 corresponding to the second type of sample pair, the total loss value L of the sample pair set can be obtained. The relevant calculation formula is L = L1 + L2.

[0136] In the training method of the Siamese network for object behavior recognition in a vehicle cabin shown in steps S1-S7, the first behavior to be recognized is set as smoking behavior, and the second behavior to be recognized is set as making a phone call behavior. However, this is only one optional embodiment provided by this application; the behavior to be recognized can also be set to other behaviors as needed. The first and second sample images in the same sample pair, or the first and third sample images in the same sample pair, can both be infrared images or both be RGB images. Images in the same sample pair set can all be infrared images or all be RGB images; multiple sample pairs can include both infrared images and RGB images.

[0137] Furthermore, this application also provides an iterative process for a deep learning model based on a Siamese neural network for infrared images. When training the Siamese network to recognize smoking and phone calls, the process may include the following: selecting 3000 positive sample infrared images of smoking in the cockpit mouth ROI and 3000 negative sample infrared images of non-smoking in the cockpit mouth ROI as training set D_1; selecting 3000 positive sample infrared images of phone calls in the cockpit eye ROI and 3000 negative sample infrared images of non-phone calls in the cockpit eye ROI as training set D_2; selecting 3000 positive sample infrared images of smoking in the full cockpit view, 3000 positive sample infrared images of phone calls in the full cockpit view, and 3000 negative sample infrared images of non-smoking and non-phone calls in the full cockpit view as training set D_3; using D_1, D_2, and D_3 as the initial training sets, training an initial multi-task Siamese neural network deep learning model based on the fusion of ROI image classification and infrared full-image target detection.

[0138] Furthermore, a multi-task Siamese neural network deep learning initial model based on the fusion of ROI image classification and infrared full-image target detection can be used to screen a dataset containing 1 million images, specifically including the following:

[0139] When screening positive samples of smoking in ROI, a multi-task Siamese neural network deep learning initial model based on the fusion of ROI image classification and infrared full-image target detection is used to output samples with a probability score of less than 0.1 of positive samples of smoking in mouth ROI and add them to the positive sample set in dataset D_1.

[0140] When screening positive samples of ROI making phone calls, a multi-task Siamese neural network deep learning initial model based on ROI image classification and infrared full-image target detection is used. Samples with a probability score of less than 0.1 for positive samples of eye ROI making phone calls are added to the positive sample set in dataset D_2.

[0141] When filtering positive samples of smoking in the whole image, a multi-task Siamese neural network deep learning initial model based on ROI image classification and infrared whole image target detection is used to output samples with a probability score of less than 0.1 of positive samples of smoking in the whole image and add them to the positive sample set in dataset D_3;

[0142] When selecting positive samples for making phone calls in the full image, a multi-task Siamese neural network deep learning initial model based on ROI image classification and infrared full image target detection is used to output samples with a probability score of less than 0.1 for positive samples making phone calls in the full image and add them to the positive sample set in dataset D_3.

[0143] When screening non-smoking negative samples of the mouth ROI, a multi-task Siamese neural network deep learning initial model based on ROI image classification and infrared full-image target detection is used to output samples with a probability score greater than 0.9 of non-smoking negative samples of the mouth ROI and add them to the negative sample set in dataset D_1.

[0144] When filtering negative samples of eye ROI without phone calls, a multi-task Siamese neural network deep learning initial model based on ROI image classification and infrared full-image target detection is used to add samples with a probability score greater than 0.9 of eye ROI without phone calls to the negative sample set in dataset D_2.

[0145] When filtering negative samples of non-smoking and non-phone-calling in the whole image, a multi-task Siamese neural network deep learning initial model based on ROI image classification and infrared whole image target detection is used to add samples with a probability score greater than 0.9 of the negative samples of non-smoking and non-phone-calling in the whole image to the negative sample set in dataset D_3.

[0146] Before adding the selected ROIs and positive and negative samples from the whole image to the initial training set, it is necessary to browse the samples to be added to check whether there are negative samples mixed in with positive samples, and whether there are positive samples mixed in with negative samples.

[0147] When the model outputs a positive sample threshold of 0.1 and a negative sample threshold of 0.9 to filter the sample set, if the number of positive and negative samples is less than 1000, then the positive sample threshold is set to 0.2 and the negative sample threshold to 0.8. This process is repeated, using thresholds of 0.1, 0.2, 0.3, and 0.4 to filter positive samples and thresholds of 0.9, 0.8, 0.7, and 0.6 to filter negative samples. Before updating the training set, the positive and negative samples to be added need to be reviewed to see if there are any positive samples mixed with negative samples, or vice versa. By continuously adding positive and negative samples to the initial training set in this way, a multi-task Siamese neural network deep learning model based on the fusion of ROI image classification and infrared full-image target detection is obtained.

[0148] In existing technologies, such as using a vehicle's cockpit perception system to detect whether a driver is smoking or making a phone call, the detection process is as follows: Face detection is performed on the driver's seat image frame sequence. If a face is detected, facial landmark regression is performed within the face's Region of Interest (ROI). Based on the landmarks, the ROI near the mouth is cropped and input into a smoking image classification model to determine if the driver is smoking. Similarly, the ROI near the eyes is cropped and input into a phone call image classification model to determine if the driver is making a phone call. The smoking image classification model determines whether the driver is smoking, and the phone call classification model determines whether the driver is making a phone call. In this process, the same classification model is used for both smoking and phone calls, but two such models are required for deployment, thus wasting valuable platform computing resources.

[0149] Based on this, this application provides an alternative implementation method in which the driver's seat image frame sequence is input into a face detection model. If a face is detected, facial key point regression is performed in the face ROI region. Based on the key points, the ROI regions near the mouth and the ROI regions near the eyes are simultaneously extracted into a multi-task image classification model to determine the behavior of smoking and making phone calls in the driver's seat.

[0150] This method requires using positive samples of smoking from the ROI near the mouth, negative samples of making phone calls from the ROI near the eyes, and negative samples (including a mixture of negative samples of smoking and negative samples from the ROI near the mouth and eyes) as a training set during the training of the multi-task infrared image classification model. The multi-task infrared image model can simultaneously output judgments for both making phone calls and smoking, saving on-board GPU computing power and thus supporting the implementation of more DMS perception tasks. Therefore, this application utilizes a multi-task approach to combine multiple functional tasks into a single model, reducing the amount of model data.

[0151] Moreover, the existing detection process only uses the ROI region to determine whether someone is smoking or making a phone call, but it does not correlate the global features. For example, when someone is smoking in the driver's seat, there is a behavioral feature of arm movement, which is different from the normal driving state. However, the ROI infrared image only involves the mouth area and lacks the arm feature, which leads to some short cigarette butts not being accurately identified.

[0152] Based on this, this application provides an alternative implementation method in which the full-image infrared image features are combined to assist the ROI infrared image in determining whether there is smoking or phone call behavior in the driver's seat.

[0153] When training the Siamese network in this application, the input sample images to be trained are two types of ROI images and a global image. When the ROI image is a mouth region image, the corresponding global image can include the driver's arm region. When the ROI image is an eye region image, the corresponding global image can also include the driver's arm region. This is because some drivers hold the phone to their eyes, i.e., the temple region, when making a phone call. Therefore, based on the training method of the Siamese network provided in this application, the trained Siamese network for object behavior recognition in the vehicle cabin can be used to determine whether the driver's behavior includes smoking or making a phone call.

[0154] Similarly, the method described above for determining whether a driver is smoking or making a phone call can also be used to determine whether a specific occupant in the vehicle is smoking or making a phone call. In other words, it can also support a wider range of OMS (On-Site Management System) sensing tasks.

[0155] Among them, DMS stands for Driver Monitoring System, which monitors the driver and can detect driver fatigue and dangerous driving behaviors. OMS stands for Occupancy Monitoring System, which monitors passengers and can detect passenger age, condition, mood, and behavior.

[0156] Based on this, this application can also provide a method for recognizing object behavior in a vehicle cabin using a Siamese network trained by this application, specifically including: acquiring a first image of a target object (such as a driver or a member) in the cabin, and determining a second image and a third object based on the first image; the second image is a local image corresponding to a first region in the first image, the first region being a region related to a first behavior to be recognized of the target object; the third image is a local image corresponding to a second region in the first image, the second region being a region related to a second behavior to be recognized of the target object; inputting the first image and the second image into a first sub-Siamese network in the Siamese network to obtain a first similarity probability between the first image feature corresponding to the first image and the second image feature corresponding to the second image, and inputting the first image and the third image into a second sub-Siamese network in the Siamese network to obtain a second similarity probability between the third image feature corresponding to the first image and the fourth image feature corresponding to the third image; determining the recognition result of the first behavior to be recognized of the target object based on the first similarity probability, and determining the recognition result of the second behavior to be recognized of the target object based on the second similarity probability; wherein, the Siamese network is a Siamese network trained based on the training method of the Siamese network provided by this application.

[0157] Specifically, the identification result of a first behavior to be identified for a target object is determined based on a first similarity probability, and the identification result of a second behavior to be identified for a target object is determined based on a second similarity probability, including: if the first similarity probability is greater than a first preset threshold, the identification result of the first behavior to be identified for the target object is determined to be the occurrence of the first behavior to be identified; and if the second similarity probability is greater than a second preset threshold, the identification result of the second behavior to be identified for the target object is determined to be the occurrence of the second behavior to be identified; wherein the first preset threshold and the second preset threshold are the same or different.

[0158] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0159] Based on the same inventive concept, this application also provides a Siamese network training apparatus for implementing the above-described Siamese network training method. The solution provided by this apparatus is similar to the implementation described in the above-described method. Therefore, the specific limitations of one or more Siamese network training apparatus embodiments provided below can be found in the limitations of the Siamese network training method described above, and will not be repeated here.

[0160] In one exemplary embodiment, such as Figure 6 As shown, a training device 300 for a Siamese network is provided, comprising: a sample acquisition module 81, a sample processing module 82, and a parameter adjustment module 83, wherein:

[0161] The sample acquisition module 81 is used to acquire a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of the sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified;

[0162] The sample processing module 82 is used to input sample pairs of different types into the corresponding sub-Twin networks in the Siamese network to be trained, obtain the first sample image features and the second sample image features corresponding to the two types of sample images in the sample pair based on the feature extraction network in the sub-Twin network, and input the first sample image features and the second sample image features into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior, respectively; wherein, different types of sample pairs correspond to different sub-Twin networks; determine the first loss value based on the distance similarity and cosine similarity between the first sample image features and the second sample image features; obtain the second loss value based on the first prediction behavior and the actual behavior corresponding to the sample pair; and obtain the third loss value based on the second prediction behavior and the actual behavior.

[0163] The parameter adjustment module 83 is used to determine the sub-loss value corresponding to each sample pair based on the first loss value, the second loss value, and the third loss value; to determine the total loss value of the sample pair set based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs respectively; and to update the network parameters of the Siamese network to be trained if the total loss value does not meet the stopping training condition, and to train the Siamese network to be trained again until the stopping training condition is met, so as to obtain the trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0164] In an exemplary embodiment, when the sample processing module 82 determines the first loss value based on the distance similarity and cosine similarity between the first sample image features and the second sample image features, it can specifically be used to: determine the sample image feature similarity based on the product of the distance similarity and cosine similarity between the first sample image features and the second sample image features; convert the sample image feature similarity into a sample image feature similarity probability within a preset value range; and determine the first loss value based on the sample image feature similarity probability, the actual behavior type value of the sample pair, and the preset loss weight.

[0165] In an exemplary embodiment, when the sample processing module 82 inputs the first sample image features and the second sample image features into the classification network in the sub-Siamese network to obtain the first prediction behavior and the second prediction behavior, it is specifically used to: input the first sample image features into the first classification network in the sub-Siamese network to obtain the first prediction behavior; and input the second sample image features into the second classification network in the sub-Siamese network to obtain the second prediction behavior.

[0166] In an exemplary embodiment, when the sample processing module 82 inputs sample pairs of various types into the corresponding sub-Siamese networks in the Siamese network to be trained, and obtains the first sample image features and the second sample image features corresponding to the two sample images in the sample pair based on the feature extraction network in the sub-Siamese network, the module specifically performs the following: for each type of sample pair, inputs one sample image from the sample pair of the type into the first feature extraction network in the sub-Siamese network to obtain the corresponding first sample image feature; and inputs the other sample image from the sample pair of the type into the second feature extraction network in the sub-Siamese network to obtain the corresponding second sample image feature.

[0167] In an exemplary embodiment, the first feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network; the second feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network.

[0168] In an exemplary embodiment, each type of sample pair includes a positive sample pair and a negative sample pair; the two sample images contained in the positive sample pair are images with the behavior to be identified; the two sample images contained in the negative sample pair are images without the behavior to be identified.

[0169] In an exemplary embodiment, the sample acquisition module 81 is used to acquire a set of sample pairs, including: acquiring multiple sample images including sample objects; determining candidate positive sample images and candidate negative sample images among the multiple sample images; for the same behavior to be identified, acquiring a first similarity between the behavior in each candidate positive sample image and the behavior to be identified; and filtering positive sample images from the multiple candidate positive sample images based on a comparison result of the first similarity and a first similarity threshold; acquiring a second similarity between the behavior in each candidate negative sample image for the same behavior and the behavior to be identified; and filtering negative sample images from the multiple candidate negative sample images based on a comparison result of the second similarity and a second similarity threshold. After image filtering is completed, from each filtered positive sample image and each negative sample image, a second sample image corresponding to the first target region and a third sample image corresponding to the second target region are obtained respectively. A first type of positive sample pair is formed based on the positive sample image and the corresponding second sample image, and a second type of positive sample pair is formed based on the positive sample image and the corresponding third sample image. A first type of negative sample pair is formed based on the negative sample image and the corresponding second sample image, and a second type of negative sample pair is formed based on the negative sample image and the corresponding third sample image. A sample pair set is obtained based on the first type of positive sample pair, the first type of negative sample pair, the second type of positive sample pair, and the second type of negative sample pair.

[0170] The modules in the aforementioned twin network training device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0171] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a twin network training method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0172] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0173] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0174] Obtain a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of the sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified;

[0175] Each type of sample pair is input into the corresponding sub-Twin network in the Siamese network to be trained. Based on the feature extraction network in the sub-Twin network, the first sample image features and the second sample image features corresponding to the two types of sample images in the sample pair are obtained respectively. The first sample image features and the second sample image features are then input into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior respectively. Different types of sample pairs correspond to different sub-Twin networks.

[0176] A first loss value is determined based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image; a second loss value is obtained based on the first predicted behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second predicted behavior and the actual behavior.

[0177] Based on the first loss value, the second loss value, and the third loss value, the sub-loss value corresponding to each sample pair is determined. Based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs, the total loss value of the sample pair set is determined. If the total loss value does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated, and the Siamese network to be trained is trained again until the stopping training condition is met, resulting in a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0178] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0179] Obtain a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of the sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified;

[0180] Each type of sample pair is input into the corresponding sub-Twin network in the Siamese network to be trained. Based on the feature extraction network in the sub-Twin network, the first sample image features and the second sample image features corresponding to the two types of sample images in the sample pair are obtained respectively. The first sample image features and the second sample image features are then input into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior respectively. Different types of sample pairs correspond to different sub-Twin networks.

[0181] A first loss value is determined based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image; a second loss value is obtained based on the first predicted behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second predicted behavior and the actual behavior.

[0182] Based on the first loss value, the second loss value, and the third loss value, the sub-loss value corresponding to each sample pair is determined. Based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs, the total loss value of the sample pair set is determined. If the total loss value does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated, and the Siamese network to be trained is trained again until the stopping training condition is met, resulting in a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0183] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0184] Obtain a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of the sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified;

[0185] Each type of sample pair is input into the corresponding sub-Twin network in the Siamese network to be trained. Based on the feature extraction network in the sub-Twin network, the first sample image features and the second sample image features corresponding to the two types of sample images in the sample pair are obtained respectively. The first sample image features and the second sample image features are then input into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior respectively. Different types of sample pairs correspond to different sub-Twin networks.

[0186] A first loss value is determined based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image; a second loss value is obtained based on the first predicted behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second predicted behavior and the actual behavior.

[0187] Based on the first loss value, the second loss value, and the third loss value, the sub-loss value corresponding to each sample pair is determined. Based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs, the total loss value of the sample pair set is determined. If the total loss value does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated, and the Siamese network to be trained is trained again until the stopping training condition is met, resulting in a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

[0188] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0189] It should also be noted that the images of user objects involved in this application are all images authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant images must comply with relevant regulations.

[0190] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0191] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0192] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for training a Siamese network, characterized in that, The method includes: Obtain a sample pair set; the sample pair set includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of a sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified; Each type of sample pair is input into the corresponding sub-Twin network in the Siamese network to be trained. Based on the feature extraction network in the sub-Twin network, the first sample image features and the second sample image features corresponding to the two sample images in the sample pair are obtained respectively. The first sample image features and the second sample image features are then input into the classification network in the sub-Twin network to obtain the first prediction behavior and the second prediction behavior respectively. Different types of sample pairs correspond to different sub-Twin networks. A first loss value is determined based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image; a second loss value is obtained based on the first predicted behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second predicted behavior and the actual behavior. Based on the first loss value, the second loss value, and the third loss value, a sub-loss value corresponding to each sample pair is determined. Based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs, the total loss value of the sample pair set is determined. If the total loss value does not meet the stopping training condition, the network parameters of the Siamese network to be trained are updated, and the Siamese network to be trained is trained again until the stopping training condition is met, resulting in a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

2. The method according to claim 1, characterized in that, The step of determining the first loss value based on the distance similarity and cosine similarity between the features of the first sample image and the features of the second sample image includes: The similarity of sample image features is determined based on the product of the distance similarity and the cosine similarity between the features of the first sample image and the features of the second sample image. The sample image feature similarity is converted into a sample image feature similarity probability within a preset value range; The first loss value is determined based on the similarity probability of the sample image features, the actual behavior type value of the sample pair, and the preset loss weight.

3. The method according to claim 1, characterized in that, The step of inputting the first sample image features and the second sample image features into the classification network in the sub-Siamese network to obtain the first prediction behavior and the second prediction behavior, respectively, includes: The features of the first sample image are input into the first classification network in the sub-Siamese network to obtain the first prediction behavior; The second sample image features are input into the second classification network in the sub-Siamese network to obtain the second prediction behavior.

4. The method according to claim 1, characterized in that, The step of inputting the sample pairs of each type into the corresponding sub-Siamese networks in the Siamese network to be trained, and obtaining the first sample image features and the second sample image features corresponding to the two sample images in the sample pair based on the feature extraction network in the sub-Siamese network, includes: For each type of sample pair, one sample image from the sample pair of the type is input into the first feature extraction network in the sub-Siamese network to obtain the corresponding first sample image features; Another sample image from the sample pair of the aforementioned type is input into the second feature extraction network in the sub-Siamese network to obtain the corresponding second sample image features.

5. The method according to claim 4, characterized in that, The first feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network; The second feature extraction network is one of the MobileNet feature extraction network and the YOLO feature extraction network.

6. The method according to claim 1, characterized in that, Each type of sample pair includes both positive and negative sample pairs; The two sample images included in the positive sample pair are both images containing the behavior to be identified; The negative sample pairs contain two types of sample images that do not contain any behavior to be identified.

7. The method according to claim 6, characterized in that, The acquisition of the sample pair set includes: Acquire multiple sample images including the sample object, and determine the candidate positive sample image and the candidate negative sample image from the multiple sample images; For the same behavior to be identified, a first similarity is obtained between the behavior in each candidate positive sample image and the behavior to be identified. Based on the comparison result of the first similarity and the first similarity threshold, positive sample images are selected from multiple candidate positive sample images. Obtain a second similarity between the behavior in each of the candidate negative sample images for the same behavior and the behavior to be identified; and filter negative sample images from multiple candidate negative sample images based on the comparison result of the second similarity and the second similarity threshold. After image filtering is completed, a second sample image corresponding to the first target region and a third sample image corresponding to the second target region are obtained from each of the filtered positive sample images and negative sample images, respectively; a first type of positive sample pair is formed based on the positive sample image and the corresponding second sample image, a second type of positive sample pair is formed based on the positive sample image and the corresponding third sample image, a first type of negative sample pair is formed based on the negative sample image and the corresponding second sample image, and a second type of negative sample pair is formed based on the negative sample image and the corresponding third sample image; The sample pair set is obtained based on the positive sample pairs of the first type, the negative sample pairs of the first type, the positive sample pairs of the second type, and the negative sample pairs of the second type.

8. A training device for a twin network, characterized in that, The device includes: A sample acquisition module is used to acquire a set of sample pairs; the set of sample pairs includes at least a first type of sample pair and a second type of sample pair; the first type of sample pair includes a first sample image and a second sample image of a sample object, and the second type of sample pair includes a first sample image and a third sample image; the second sample image is a local image of a first target region in the first sample image, and the third sample image is a local image of a second target region in the first sample image, wherein the first target region is a region related to a first behavior to be identified, and the second target region is a region related to a second behavior to be identified; A sample processing module is used to input the sample pairs of each type into the corresponding sub-Twin networks in the Siamese network to be trained, obtain first sample image features and second sample image features corresponding to the two types of sample images in the sample pair based on the feature extraction network in the sub-Twin network, and input the first sample image features and second sample image features into the classification network in the sub-Twin network to obtain a first prediction behavior and a second prediction behavior, respectively; wherein, different types of sample pairs correspond to different sub-Twin networks; a first loss value is determined based on the distance similarity and cosine similarity between the first sample image features and the second sample image features; a second loss value is obtained based on the first prediction behavior and the actual behavior corresponding to the sample pair; and a third loss value is obtained based on the second prediction behavior and the actual behavior. The parameter adjustment module is used to determine the sub-loss value corresponding to each sample pair based on the first loss value, the second loss value, and the third loss value; to determine the total loss value of the sample pair set based on the sub-loss values ​​corresponding to at least the first type and the second type of sample pairs; and to update the network parameters of the Siamese network to be trained and retrain the Siamese network to be trained if the total loss value does not meet the stopping training condition, until the stopping training condition is met, thereby obtaining a trained Siamese network. The trained Siamese network is applied to the recognition of at least two behaviors of objects in the vehicle cabin.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Behavior recognition method and device and computer readable storage medium

    CN112329719A

  • Image recognition model training method and device, electronic equipment and storage medium

    CN117218400A