Sample generation method, image recognition model training method and corresponding device

By clustering the on-board image sets to generate triple sample pairs, the problem of training a small number of samples in on-board image recognition is solved, and the convergence stability and recognition accuracy of the model are improved.

CN118351399BActive Publication Date: 2025-05-13BYD CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410758335.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-05-13
Estimated Expiration
2044-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively use a small number of samples for model training in vehicle-mounted image recognition, and the convergence stability and recognition accuracy of the model are insufficient.

Method used

By determining multiple image sets and clustering them, positive and negative sample pairs in intra-class and inter-class clusters are generated to form triple sample pairs to train the image recognition model.

Benefits of technology

The convergence stability and recognition accuracy of the image recognition model are improved, and the invalid sample problems caused by improper sample selection are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118351399B_ABST
    Figure CN118351399B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a sample generation method, a training method for an image recognition model and a corresponding device, and to the field of data processing technology. The method includes: determining multiple image sets; for each image set, clustering multiple images corresponding to the image set to obtain multiple intra-class clusters, and selecting images in different intra-class clusters to form positive sample pairs; clustering multiple image sets to obtain multiple inter-class clusters, and selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs; generating triplet sample pairs based on multiple image sets, positive sample pairs and negative sample pairs. By selecting images with lower similarity in different intra-class clusters to form positive sample pairs, the situation of generating invalid positive sample pairs due to selecting the same image with higher similarity can be reduced. Selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs with higher similarity can reduce the situation of generating invalid negative sample pairs due to selecting different images with lower similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular, to a sample generation method, an image recognition model training method and corresponding devices. Background Art

[0002] In-vehicle image recognition mainly applies image recognition technology to vehicles. When image recognition technology is installed in a vehicle, the vehicle can scan the in-vehicle image to identify the corresponding identity information, so as to realize personalized services such as door unlocking, ignition start, seat adjustment and route planning, as well as in-vehicle payment and other functions.

[0003] In the related art, a large number of images of drivers are collected and used as training samples to train an image recognition model, and then the trained image recognition model is used to identify the driver's identity information to achieve personalized vehicle services and functions. Summary of the invention

[0004] The purpose of the present disclosure is to provide a sample generation method, an image recognition model training method and corresponding devices to solve the technical problems existing in the above-mentioned related technologies.

[0005] In order to achieve the above objectives, in a first aspect, the present disclosure provides a sample generation method, comprising:

[0006] Determine a plurality of image sets, each of the image sets comprising a plurality of images corresponding to the same identity information;

[0007] For each of the image sets, clustering the multiple images corresponding to the image set to obtain multiple intra-class clusters, selecting images from different intra-class clusters to form positive sample pairs, wherein the similarity between the features corresponding to different images in the same intra-class cluster is less than a first similarity threshold;

[0008] Clustering the multiple image sets to obtain multiple inter-class clusters, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs, wherein the similarity between the features corresponding to the different image sets in the same inter-class cluster is less than a second similarity threshold;

[0009] A triplet of sample pairs is generated according to the multiple image sets, the positive sample pairs, and the negative sample pairs.

[0010] Optionally, clustering the multiple images corresponding to the image set to obtain multiple intra-class clusters includes:

[0011] Calculate the similarity between every two images in the image set;

[0012] Considering the images in the image set as nodes, connecting the nodes corresponding to the two images whose similarity is less than the first similarity threshold, to obtain a first undirected graph;

[0013] The image corresponding to each subgraph in the first undirected graph is determined as an intra-cluster of a class, so as to obtain a plurality of intra-cluster of the class.

[0014] Optionally, selecting images from different in-class clusters to form positive sample pairs includes:

[0015] Randomly selecting a first image from a first in-class cluster and randomly selecting a second image from a second in-class cluster, wherein the first in-class cluster and the second in-class cluster are different in-class clusters among the multiple in-class clusters;

[0016] The first image and the second image constitute a positive sample pair.

[0017] Optionally, clustering the multiple image sets to obtain multiple inter-class clusters includes:

[0018] Calculating the similarity between the central features of every two image sets in the plurality of image sets, the central feature of the image being the image in the image set that meets a preset condition;

[0019] The central features of the image sets are regarded as nodes, and the nodes corresponding to the central features of the two image sets whose similarities are less than the second similarity threshold are connected to obtain a second undirected graph;

[0020] The image set corresponding to each subgraph in the second undirected graph is determined as an inter-class cluster to obtain the multiple inter-class clusters.

[0021] Optionally, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs includes:

[0022] Randomly selecting a third image and a fourth image from the same between-class cluster, wherein the third image and the fourth image are images corresponding to different image sets;

[0023] The third image and the fourth image form a negative sample pair.

[0024] Optionally, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs includes:

[0025] For each image set in the same between-class cluster, randomly selecting a fifth image from the image set;

[0026] Randomly select a sixth image from each of the multiple intra-class clusters of the remaining image set, wherein the remaining image set is the image set remaining after excluding the image set in the same inter-class cluster;

[0027] The fifth image and the sixth image constitute the negative sample pair.

[0028] Optionally, determining a plurality of image sets includes:

[0029] A plurality of image sets captured by the vehicle-mounted image capture device are determined.

[0030] In a second aspect, the present disclosure provides a method for training an image recognition model, the method comprising:

[0031] The first image recognition model is trained according to the triplet sample pairs generated by the sample generation method according to any one of the first aspects of the present disclosure.

[0032] Optionally, before training the first image recognition model, the method further includes:

[0033] A vehicle-mounted image set in a vehicle-mounted image acquisition device is obtained, and a second image recognition model is trained according to the vehicle-mounted image set by using a normalized exponential softmax loss function to obtain the first image recognition model.

[0034] Optionally, before training the second image recognition model by using a normalized exponential softmax loss function to obtain the first image recognition model, the method further includes:

[0035] Get RGB image set;

[0036] Performing grayscale processing on the RGB image set to obtain a grayscale image set;

[0037] The third image recognition model is pre-trained according to the grayscale image set to obtain the second image recognition model.

[0038] In a third aspect, the present disclosure provides a non-temporary computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods provided in the first and second aspects of the present disclosure.

[0039] In a fourth aspect, the present disclosure provides a controller, comprising:

[0040] a memory having a computer program stored thereon;

[0041] A processor is used to execute the computer program in the memory to implement the steps of any one of the methods provided in the first aspect and the second aspect of the present disclosure.

[0042] In a fifth aspect, the present disclosure provides a vehicle, comprising the controller provided in the fourth aspect of the present disclosure.

[0043] In a sixth aspect, the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any one of the methods provided in the first and second aspects of the present disclosure.

[0044] Through the above technical solution, the similarity of multiple images corresponding to the same intra-class cluster is high, and the similarity of images corresponding to different intra-class clusters is low. Selecting images with lower similarity in different intra-class clusters to form positive sample pairs can reduce the situation where invalid positive sample pairs are generated due to selecting the same image with higher similarity. In addition, the similarity of images of multiple image sets corresponding to the same inter-class cluster is high. Selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs with higher similarity can reduce the situation where invalid negative sample pairs are generated due to selecting different images with lower similarity. Therefore, a large number of triple sample pairs can be generated based on the positive sample pairs, negative sample pairs, and image sets. Furthermore, when the triple sample pairs are used to train the first image recognition model, the convergence stability and recognition accuracy of the first image recognition model can be improved.

[0045] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:

[0047] Figure 1 It is a flowchart of vehicle-mounted face recognition in related technologies.

[0048] Figure 2 It is a flowchart of training models in related technologies.

[0049] Figure 3 is a schematic diagram showing a sample generation method according to an exemplary embodiment of the present disclosure.

[0050] Figure 4 is a schematic diagram showing a first undirected graph according to an exemplary embodiment of the present disclosure.

[0051] Figure 5 The figure is a flowchart showing a method for training an image recognition model according to an exemplary embodiment of the present disclosure.

[0052] Figure 6 It is a schematic diagram showing a process of selecting triplet sample pairs according to an exemplary embodiment of the present disclosure.

[0053] Figure 7 is a schematic diagram showing a sample generating device according to an exemplary embodiment of the present disclosure.

[0054] Figure 8 is a block diagram showing a vehicle according to an exemplary embodiment. DETAILED DESCRIPTION

[0055] The specific implementation of the present disclosure is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the present disclosure, and is not used to limit the present disclosure.

[0056] Image recognition technology is mainly a recognition technology that analyzes and processes images through computer vision technology. However, in related technologies, when performing image recognition, it is impossible to train the image recognition model with a small number of samples, or it is impossible to obtain a target image recognition model with stable convergence. This embodiment takes face recognition technology in image recognition technology as an example for explanation as follows.

[0057] When face recognition technology is installed in vehicles, Figure 1 As shown in the figure, the vehicle-mounted face recognition technology mainly includes three processes: training model, registering identity and identifying. In the process of training the model, a large number of driver's face images are collected through the vehicle-mounted camera, and the face images are preprocessed. Then, the face recognition model is trained through the deep learning algorithm according to the preprocessed face images to obtain the target face recognition model. In the process of registering identity, the driver enters his own identity information and records his own face image through the vehicle-mounted camera, wherein the face image includes the front and multi-angle face images of the driver. In the process of identification, the real-time face image of the driver is collected through the vehicle-mounted camera, and when the quality of the real-time face image is qualified, the matching rate of the real-time face image and the recorded face image is calculated through the target face recognition model, and the identity information of the corresponding driver is obtained according to the matching rate, and returned to the terminal, so that the corresponding vehicle control can be realized.

[0058] like Figure 2 As shown in the figure, in the process of training the model, a large number of driver's face images are first collected, and then the face area image is extracted from the face image through the face detection model, and the face area image is affine transformed to align the face area image to the area at the same position as the preset face template. Then the aligned face image is sent to the face recognition model, and by designing a reasonable loss function, the face recognition model is converged to obtain the target face recognition model, which can then make the target face recognition model distinguish drivers with different identity information.

[0059] During the model training process, the loss function in the face recognition model is trained mainly through the loss functions provided in the first related technology and the second related technology.

[0060] In the first related technology, it is mainly based on the softmax (normalized exponential function) and its variant loss functions based on margin (the distance between the training samples closest to the decision boundary and the force boundary), such as CosFace (cosine face), SphereFace (spherical face), ArcFace (arc face), etc. This type of loss function regards a large number of input face images as classification tasks, and regards face images corresponding to the same identity information as the same type of face images. When face images collected from drivers with different identity information are used as training samples, the face recognition model outputs a certain number of types of face images after training. The ultimate goal of training is to ensure that a large number of input face images are correctly classified.

[0061] In the second related technology, the loss function is mainly based on metric learning, such as triplet loss, contrastive loss, etc. This type of method constructs triplet sample pairs so that the distance between face images with the same identity information is smaller than the distance between face images with different identity information. This method can combine a small number of face images with different strategies to form a large number of training samples, which is more suitable for application scenarios with fewer face images.

[0062] When the triplet loss function is used as the loss function in the face recognition model to train face images, the triplet loss can be expressed by the following calculation formula.

[0063]

[0064] in, is the input triplet sample, a is the anchor, p is the positive, and n is the negative. And a and p can represent different face images corresponding to the same identity information, and a and n can represent face images corresponding to different identity information.

[0065] In the second related technology, triple sample pairs are selected by the following method. In each round of iterative training of the face recognition model, triple sample pairs are generated online using the collected face images and the face recognition model. The specific method includes: in each round of iteration, all face images corresponding to p different identity information are randomly extracted from a large number of face images, and k face images are randomly extracted from the face images corresponding to each of the p identity information, and the k extracted face images constitute the current sample. Then, for each face image in each identity information, the face image and other face images corresponding to the identity information constitute a positive sample pair, and then a positive sample pair can be generated. positive sample pairs. Then, for each positive sample pair, the positive sample pair is combined with each face image in other identity information to form a negative sample pair, so a total of triplet sample pairs .

[0066] However, the inventors found that when using the loss function in the first related technology to train a face recognition model, a large number of face image samples are required, but the face images obtained from the vehicle are a small number of samples, and thus it is impossible to train a face recognition model on the face images obtained from the vehicle.

[0067] When the face recognition model is trained by the triplet loss function in the second related technology, the face images in the current sample are randomly extracted, and there may be face images with great differences such as men and women, which cannot form valuable negative sample pairs, resulting in the inability to converge the trained face recognition model, which is an invalid sample. Positive sample pairs are formed by randomly extracting from the face images corresponding to each identity information. The positive sample pairs may include a large number of face images with high similarity, and thus the trained face recognition model cannot converge, which is an invalid sample. And by the method of randomly extracting positive sample pairs and negative sample pairs, it may be impossible to extract positive sample pairs with low similarity and negative sample pairs with high similarity, which makes the accuracy of the trained target face recognition model uncertain, and may make the convergence of the target face recognition model unstable.

[0068] In view of this, the present disclosure provides a sample generation method, an image recognition model training method and corresponding devices to solve the problems existing in the above-mentioned related technologies.

[0069] like Figure 3 As shown, Figure 3 is a schematic diagram showing a sample generation method according to an exemplary embodiment of the present disclosure, referring to Figure 3 ,include:

[0070] S301: Determine a plurality of image sets, each of the image sets including a plurality of images corresponding to the same identity information;

[0071] S302: for each of the image sets, clustering multiple images corresponding to the image set to obtain multiple intra-class clusters, selecting images from different intra-class clusters to form positive sample pairs, wherein the similarity between features corresponding to different images in the same intra-class cluster is less than a first similarity threshold;

[0072] S303: clustering the multiple image sets to obtain multiple inter-class clusters, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs, wherein the similarity between the features corresponding to different image sets in the same inter-class cluster is less than a second similarity threshold;

[0073] S304: Generate triplet sample pairs according to the multiple image sets, the positive sample pairs, and the negative sample pairs.

[0074] Through the above technical solution, the similarity of multiple images corresponding to the same intra-class cluster is high, and the similarity of images corresponding to different intra-class clusters is low. Selecting images with lower similarity in different intra-class clusters to form positive sample pairs can reduce the situation where invalid positive sample pairs are generated due to selecting the same image with higher similarity. In addition, the similarity of images of multiple image sets corresponding to the same inter-class cluster is high. Selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs with higher similarity can reduce the situation where invalid negative sample pairs are generated due to selecting different images with lower similarity. Therefore, a large number of triple sample pairs can be generated based on the positive sample pairs, negative sample pairs, and image sets. When the triple sample pairs are used to train the first image recognition model, the convergence stability and recognition accuracy of the first image recognition model can be improved.

[0075] In addition, compared with the method of training face recognition models through a large number of face images in related technologies, the technical solution provided by the present invention does not require a large number of images for training, and can achieve a large number of triple sample pairs composed of a small number of images, and train the first image recognition model through the triple sample pairs.

[0076] In order to enable those skilled in the art to better understand the sample generation method provided by the present disclosure, the above steps are described in detail with examples below.

[0077] For example, an image set may be a plurality of images corresponding to a piece of identity information collected. Multiple image sets may be a plurality of images corresponding to multiple pieces of identity information. Among them, an image set may be an image collected by a vehicle-mounted camera, which is not limited in the embodiments of the present disclosure. An image set may be a face image set, and the multiple images included in the image set may be multiple face images, which is not specifically limited in the embodiments of the present disclosure.

[0078] For example, clustering can be to divide multiple data into several categories, so that the data in the same category are similar to each other, and the data between different categories are relatively different. And the same image set can be multiple images under the same identity information, among which there may be images with high similarity and images with low similarity. In this way, multiple images in the image set can be clustered, and images with high similarity can be divided into an intra-class cluster, so that multiple intra-class clusters can be obtained, and the similarity of images between different intra-class clusters is low. Images can be selected from different intra-class clusters to form positive sample pairs, so that the similarity of the positive sample pairs can be low, and the occurrence of positive sample pairs with high similarity can be avoided. Among them, the similarity between the features corresponding to different images in the same intra-class cluster is less than the first similarity threshold, so that the similarity of different images in the same intra-class cluster can be high. Among them, the first similarity threshold can be a critical threshold for the similarity between two images.

[0079] By clustering multiple images in the image set, the similarity of multiple images corresponding to the same cluster is higher, and the similarity of images corresponding to different clusters is lower. Images with lower similarity are selected from different clusters to form positive sample pairs, which can reduce the situation of invalid positive sample pairs caused by selecting the same image with higher similarity.

[0080] In a possible manner, clustering the multiple images corresponding to the image set to obtain multiple intra-class clusters includes:

[0081] Calculate the similarity between every two images in the image set;

[0082] Considering the images in the image set as nodes, connecting the nodes corresponding to the two images whose similarity is less than the first similarity threshold, to obtain a first undirected graph;

[0083] The image corresponding to each subgraph in the first undirected graph is determined as an intra-cluster of a class, so as to obtain a plurality of intra-cluster of the class.

[0084] It should be understood that the similarity can be the Euclidean distance, which is not specifically limited in the embodiments of the present disclosure. When the similarity is the Euclidean distance, the Euclidean distance between every two images in the image set can be expressed by the following calculation formula.

[0085]

[0086] Where N can be the feature dimension of multiple images in the image set, and m can be the index value. can be the Euclidean distance, , Both can be eigenvalues ​​in the matrix vector formed by multiple images in the image set.

[0087] When the Euclidean distance is less than the first similarity threshold, it can represent that the similarity between the two images is high, and the two images can be classified into the same intra-class cluster. When the Euclidean distance is greater than the first similarity threshold, it can represent that the similarity between the two images is low, and the two images may not belong to the same intra-class cluster. The first undirected graph can be a mathematical structure composed of vertices and edges connecting these vertices.

[0088] Then, each image in the image set can be regarded as a node, and the nodes corresponding to the two images whose Euclidean distance is less than the first similarity threshold are connected to obtain a first undirected graph. In the first undirected graph, multiple nodes connected together can be regarded as a subgraph. Then, in the first undirected graph, multiple subgraphs can be formed or one subgraph can be formed. Among them, each subgraph can include two nodes, three nodes, four nodes and other nodes of varying numbers. When the subgraph includes two nodes, it can represent that the similarity of the two images corresponding to the nodes is high. When the subgraph includes three nodes, it can represent that the similarity of the images corresponding to the three nodes is high, and so on. Then, the subgraph in the first undirected graph can be determined as an intra-class cluster, and multiple intra-class clusters can be obtained.

[0089] like Figure 4 As shown, Figure 4 It is a schematic diagram showing a first undirected graph according to an exemplary embodiment of the present disclosure. In the first undirected graph, a first subgraph connected by eight nodes ah, a second subgraph connected by three nodes ln, and a third subgraph connected by three nodes ik are included. In the first subgraph, the similarity of images corresponding to any two of the eight nodes ah is high, and the first subgraph can be determined as an inner cluster of a class. In the second subgraph, the similarity of images corresponding to the three nodes ln is high, and the second subgraph can be determined as an inner cluster of a class. In the third subgraph, the similarity of images corresponding to the three nodes ik is high, and the third subgraph can be determined as an inner cluster of a class. Then, one image is randomly selected from the first subgraph and the second subgraph, and the similarity of the two images is low. One image is randomly selected from the second subgraph and the third subgraph, and the similarity of the two images is low. One image is randomly selected from the first subgraph and the third subgraph, and the similarity of the two images is low. Thus, the three subgraphs can be determined as three inner clusters in the first undirected graph.

[0090] According to the similarity, multiple images in the image set are divided into multiple intra-class clusters, so that the similarity of images in multiple image sets corresponding to the same inter-class cluster is higher. In the same inter-class cluster, images corresponding to different image sets are selected to form negative sample pairs with higher similarity, which can reduce the situation of invalid negative sample pairs generated by selecting different images with lower similarity.

[0091] In a possible manner, selecting images from different in-class clusters to form positive sample pairs includes:

[0092] Randomly selecting a first image from a first in-class cluster and randomly selecting a second image from a second in-class cluster, wherein the first in-class cluster and the second in-class cluster are different in-class clusters among the multiple in-class clusters;

[0093] The first image and the second image constitute a positive sample pair.

[0094] It should be understood that when forming a positive sample pair, two images with low similarity in the same image set need to be formed into a positive sample pair. Thus, a first image can be randomly selected from a first-class cluster, and a second image can be randomly selected from a second-class cluster, and the first-class cluster and the second-class cluster belong to different in-class clusters among multiple in-class clusters. In other words, the first image and the second image are images with low similarity, and then the first image and the second image can be formed into a positive sample pair, which is a positive sample pair with low similarity. This can avoid selecting negative sample pairs with high similarity. Thus, when the image set includes M in-class clusters, it can be formed. Positive sample pairs.

[0095] Selecting images with lower similarity from clusters within different classes to form positive sample pairs can reduce the situation where invalid positive sample pairs are generated due to selecting the same image with higher similarity. And when the positive sample pairs are used for subsequent training of the first image recognition model, the situation where invalid training occurs can be reduced.

[0096] For example, an inter-class cluster can be a set of multiple images corresponding to different identity information with high similarity. In multiple image sets, there may be images with different identity information but similarity. Therefore, multiple image sets can be clustered, and multiple image sets with high similarity can be divided into an inter-class cluster, and multiple inter-class clusters can be obtained. After that, images corresponding to different image sets can be selected in the same inter-class cluster to form negative sample pairs. In this way, negative sample pairs with high similarity can be formed, and negative sample pairs with low similarity can be avoided.

[0097] By dividing multiple image sets into multiple inter-class clusters through a clustering method, and selecting images corresponding to different image sets in each inter-class cluster to form negative sample pairs, negative sample pairs with high similarity can be obtained. The situation of generating invalid negative sample pairs due to selecting different images with low similarity can be reduced. Furthermore, when the negative sample pairs are used to train the first image recognition model, the situation of invalid training can be reduced.

[0098] In a possible manner, clustering the multiple image sets to obtain multiple inter-class clusters includes:

[0099] Calculating the similarity between the central features of every two image sets in the plurality of image sets, the central feature of the image being the image in the image set that meets a preset condition;

[0100] The central features of the image sets are regarded as nodes, and the nodes corresponding to the central features of the two image sets whose similarities are less than the second similarity threshold are connected to obtain a second undirected graph;

[0101] The image set corresponding to each subgraph in the second undirected graph is determined as an inter-class cluster to obtain the multiple inter-class clusters.

[0102] It should be understood that the central feature of the image set may be the images in the image set that meet the preset conditions.

[0103] When determining the central feature of an image set, the sum of the distances between each image in the image set and the remaining images can be calculated, and the image corresponding to the minimum distance sum can be determined as the central feature of the image set. Then the preset condition can be that the sum of the distances between the images in the image set and the remaining images reaches the minimum value. The central feature of the image set can be calculated by the following calculation formula.

[0104]

[0105]

[0106] Among them, K is the number of images in the image set, C is the central feature of the image set, is the Euclidean distance between the features corresponding to two images in the image set, can be the distance and, Can be the minimum distance and.

[0107] In this embodiment, the similarity may be a Euclidean distance, which is not specifically limited in the embodiment of the present disclosure. When the similarity is a Euclidean distance, the Euclidean distance between the central features of every two image sets in the plurality of image sets may be calculated based on the central features of the image sets. The Euclidean distance may be expressed by the following calculation formula.

[0108]

[0109] in, is the Euclidean distance between the central features of each two image sets, N is the feature dimension of the central feature of the image set, and m is the index value. , are the central features of the image set.

[0110] When the Euclidean distance is less than the second similarity threshold, it can represent that the images corresponding to the two image sets corresponding to the similarity are highly similar; otherwise, it can represent that the images corresponding to the two image sets corresponding to the similarity are relatively low similarity. The central feature corresponding to each image set can be regarded as a node, and the nodes corresponding to the central features of the two image sets whose Euclidean distance is less than the second similarity threshold are connected to obtain a second undirected graph. In the second undirected graph, one subgraph or multiple subgraphs can be included. In each subgraph, one image set corresponding to a node can be included, or image sets corresponding to multiple nodes can be included. In each subgraph, the images between different image sets are highly similar. In other words, each subgraph can be an image corresponding to different identity information with high similarity. Each subgraph is determined as an inter-class cluster, and then multiple inter-class clusters can be obtained according to the second undirected graph.

[0111] By clustering the central features of multiple image sets, the multiple image sets are divided into multiple inter-class clusters, and the inter-class clusters can be used in the subsequent determination of negative sample pairs, thereby avoiding the occurrence of negative sample pairs with low similarity.

[0112] In a possible manner, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs includes:

[0113] Randomly selecting a third image and a fourth image from the same between-class cluster, wherein the third image and the fourth image are images corresponding to different image sets;

[0114] The third image and the fourth image form a negative sample pair.

[0115] It should be understood that in each inter-class cluster, multiple image sets with high similarity can be included. The third image and the fourth image can be selected from the same inter-class cluster, and the third image and the fourth image belong to images corresponding to different image sets. In other words, the third image and the fourth image belong to images with different identity information but high similarity. Furthermore, the third image and the fourth image can be formed into a negative sample pair, so that the similarity of the negative sample pair is high, and the occurrence of negative sample pairs with low similarity can be reduced.

[0116] By selecting negative sample pairs with higher similarity in the same inter-class cluster, the situation of generating invalid negative sample pairs due to selecting different images with lower similarity can be reduced. Furthermore, when the negative sample pairs are used for subsequent training of the first image recognition model, the situation of invalid training can be reduced, and the training accuracy of the first image recognition model can be improved.

[0117] In a possible manner, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs includes:

[0118] For each image set in the same between-class cluster, randomly selecting a fifth image from the image set;

[0119] Randomly select a sixth image from each of the multiple intra-class clusters of the remaining image set, wherein the remaining image set is the image set remaining after excluding the image set in the same inter-class cluster;

[0120] The fifth image and the sixth image constitute the negative sample pair.

[0121] It should be understood that for the same inter-class cluster, the fifth image can be randomly selected from each image set in the inter-class cluster. Then, multiple intra-class clusters in the remaining image set can be determined, and the sixth image is randomly selected from each intra-class cluster of the multiple intra-class clusters. The remaining image set is the image set remaining after the image set is removed from the same inter-class cluster. The fifth image set and the sixth image set constitute a negative sample pair, and a negative sample pair with a high similarity can be obtained.

[0122] For example, when the inter-class cluster includes four image sets from the first image set to the fourth image set, when calculating the negative sample pairs formed by the first image set, a fifth image is randomly selected from the first image set. When the second image set includes 4 intra-class clusters, the third image set includes 5 intra-class clusters, and the fourth image set includes 8 intra-class clusters, it can be calculated that the total number of intra-class clusters included in the second image set to the fourth image set is 17. A sixth image can be randomly selected from each of the 17 intra-class clusters, and 17 sixth images can be obtained. The fifth image can be respectively combined with the 17 sixth images to form 17 negative sample pairs. The method for calculating the negative sample pairs formed from the second image set to the fourth image set is the same as that formed by the first image set, and will not be repeated here.

[0123] By combining the images corresponding to each image set in the same inter-class cluster with the images corresponding to each intra-class cluster in the remaining image set as negative sample pairs, the selected negative sample pairs can be negative sample pairs with higher similarity, thereby avoiding the occurrence of negative sample pairs with lower similarity. And by selecting valuable negative sample pairs based on the similarity between images, and training the first image recognition model based on the negative sample pairs, the problem of oscillation of the first image recognition model caused by randomly selecting sample pairs can be avoided, and the occurrence of invalid training can be reduced.

[0124] For example, a triplet sample pair may be generated based on multiple image sets, positive sample pairs, and negative sample pairs. When generating the triplet sample pair, each positive sample pair in each image set may be combined with a negative sample pair formed by the inter-class cluster of the image set into a triplet sample pair.

[0125] For example, when the image set includes A intra-class clusters, the image set can form positive sample pairs. When the inter-class cluster of the image set includes 5 image sets, the intra-class clusters of each image set except the current image set include P1, P2, P3 and P4 intra-class clusters respectively. Randomly extract an image from each of these intra-class clusters, and a total of P1+P2+P3+P4 images can be extracted. Each positive sample pair in the image set is combined with the multiple images, and a total of For each image set, the above method is used to calculate the number of triple sample pairs, and a large number of triple sample pairs can be formed from a small number of images.

[0126] When the positive sample pairs with lower similarity and the negative sample pairs with higher similarity are combined into a triple sample pair, and the triple sample pair is used for subsequent training of the first image recognition model, the situation where the positive sample pairs with lower similarity and the negative sample pairs with higher similarity cannot be selected when the triple sample pairs are randomly selected can be reduced. The problem of oscillation in the training of the first image recognition model caused by the random selection of positive sample pairs and negative sample pairs can also be reduced, and the accuracy of the trained first image recognition model can be made more accurate, thereby making the convergence stability of the first image recognition model better.

[0127] In a possible manner, the determining of the plurality of image sets includes:

[0128] A plurality of image sets captured by the vehicle-mounted image capture device are determined.

[0129] It should be understood that the vehicle-mounted image acquisition device may be a vehicle-mounted camera for acquiring images of the vehicle driver. The vehicle-mounted camera may be an in-vehicle monitoring camera, a panoramic camera, or other cameras, which are not specifically limited in the embodiments of the present disclosure. A small number of vehicle-mounted image sets may be acquired through the vehicle-mounted image acquisition device, and then feature data in the vehicle-mounted image sets may be extracted through an image recognition model to obtain multiple image sets.

[0130] Based on the same concept, this embodiment also discloses a method for training an image recognition model, the method comprising:

[0131] The first image recognition model is trained based on the triplet sample pairs generated by the sample generation method disclosed in this embodiment.

[0132] For example, the sample production method disclosed in this embodiment can generate a large number of triple sample pairs based on a small number of image sets, and use the large number of triple sample pairs to train the first image recognition model. The first image recognition model can be a face recognition model, which is not specifically limited in the embodiment disclosed in this disclosure. Compared with the related art of obtaining a large number of face images to train the face recognition model, the training method of the first image recognition model provided in this embodiment does not need to provide a large number of images, and the first image recognition model can be trained to obtain a target image recognition model.

[0133] In a possible manner, before training the first image recognition model, the method further includes:

[0134] A vehicle-mounted image set in a vehicle-mounted image acquisition device is obtained, and a second image recognition model is trained according to the vehicle-mounted image set by using a normalized exponential softmax loss function to obtain the first image recognition model.

[0135] It should be understood that a vehicle-mounted image set collected by a vehicle-mounted image acquisition device is obtained, wherein the vehicle-mounted image acquisition device may be a vehicle-mounted camera, and the vehicle-mounted camera may be an in-vehicle monitoring camera, a panoramic camera, or other cameras, and the embodiments of the present disclosure do not specifically limit this. Afterwards, the second image recognition model may be trained by the vehicle-mounted image set through a normalized exponential softmax loss function to obtain a first image recognition model. Afterwards, the first image recognition model is trained by triplet sample pairs. Among them, the training process of training the second image recognition model by a normalized exponential softmax loss function is a training process in the related art, and will not be repeated here.

[0136] First, the second image recognition model is trained by the normalized exponential softmax loss function, and then the first image recognition model is trained by the triple sample pairs, which can avoid the overfitting phenomenon that occurs when the second image recognition model is trained only by the softmax loss function, and can avoid the low generalization of the first image recognition model. The triple sample pairs can expand a small amount of vehicle images into a large amount of triple sample pair data, thereby avoiding the problem of being unable to train due to the small number of images in the vehicle image set, and can reduce the difficulty of sample acquisition.

[0137] The following operations are performed during each iteration: extract multiple image sets from the vehicle image set through the first image recognition model, train the first image recognition model according to the multiple image sets through the training method of the above-mentioned image recognition model, obtain a new first image recognition model and a loss value, and judge whether the loss value is less than the preset threshold. If the loss value is greater than the preset threshold, extract multiple new image sets from the vehicle image set through the new first image recognition model until the loss value obtained is less than the preset threshold, and stop the iteration. The first image recognition model obtained from the last iteration is used as the target image recognition model, and the target image recognition model is used to identify images in the vehicle. Among them, the preset threshold can represent the threshold at which the model reaches the convergence condition.

[0138] By extracting multiple image sets from the vehicle-mounted image set through the second image recognition model, and grouping the multiple image sets into triplet sample pairs, and performing iterative optimization training, the accuracy of the target image recognition model can be improved and the convergence stability of the target image recognition model can be improved.

[0139] In a possible manner, before training the second image recognition model by using a normalized exponential softmax loss function to obtain the first image recognition model, the method further includes:

[0140] Get RGB image set;

[0141] Performing grayscale processing on the RGB image set to obtain a grayscale image set;

[0142] The third image recognition model is pre-trained according to the grayscale image set to obtain the second image recognition model.

[0143] It should be understood that the RGB image set may be a large number of various image sets directly downloaded from the Internet. The grayscale processing of the RGB image set may be performed by the following calculation formula to obtain a grayscale image set.

[0144]

[0145] in, For the converted The pixel value of the point, is the pixel value of the pixel point in the RGB image set, is the R channel point before transformation The pixel value of G channel point before conversion The pixel value of is the B channel point before conversion The pixel value of .

[0146] Afterwards, the third image recognition model is pre-trained with a grayscale image set to obtain a second image recognition model, wherein the specific training process of pre-training the third image recognition model with a grayscale image set is the training process in the related art and will not be described here. The second image recognition model can then be applied to subsequent fine-tuning with an on-board image set to obtain a first image recognition model. By pre-training the third image recognition model with a large number of grayscale image sets after grayscale processing from a large number of RGB images, the convergence speed and generalization of the first image recognition model can be improved.

[0147] Reference Figure 5 , Figure 5 is a flowchart showing a method for training an image recognition model according to an exemplary embodiment of the present disclosure, such as Figure 5 As shown, the process of the image recognition model includes the following steps.

[0148] S500: Train a third image recognition model using an RGB image set to obtain a second image recognition model.

[0149] S501: Train the second image recognition model through a normalized exponential softmax loss function to obtain a first image recognition model.

[0150] S502: Generate triplet sample pairs.

[0151] S503: Train a first image recognition model.

[0152] S504: Iterate optimization and return to step S502.

[0153] S505: Determine whether the loss value reaches a preset threshold value, if so, execute step S506, otherwise execute step S504.

[0154] S506: Obtain a target image recognition model.

[0155] Reference Figure 6 , Figure 6 FIG. 1 is a schematic diagram showing a process of selecting a triplet sample pair according to an exemplary embodiment of the present disclosure. Figure 6 As shown, the process of the method for selecting triplet sample pairs includes the following steps.

[0156] S600: extract multiple image sets through the first image recognition model, and execute step S601 and step S603 respectively.

[0157] S601: Clustering multiple images corresponding to each image set.

[0158] S602: Select positive sample pairs within clusters in different classes and execute step S606.

[0159] S603: Calculate the central feature of each image set.

[0160] S604: Clustering the central features of all image sets.

[0161] S605: Select negative sample pairs within the same inter-class cluster and execute step S606.

[0162] S606: Construct triplet sample pairs.

[0163] The specific implementation methods of the above-mentioned process steps have been illustrated in detail above and will not be repeated here. In addition, it should be understood that for the above-mentioned system embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited to the order of actions described above. Secondly, those skilled in the art should also know that the embodiments described above belong to preferred embodiments, and the steps involved are not necessarily required by the present disclosure.

[0164] Through the above technical solution, by selecting images with lower similarity in different intra-class clusters to form positive sample pairs, the situation of generating invalid positive sample pairs due to selecting the same image with higher similarity can be reduced. By selecting different images in the same inter-class cluster to form negative sample pairs, the situation of generating invalid negative sample pairs due to selecting different images with lower similarity can be reduced. Thus, according to the positive sample pairs, negative sample pairs, and image sets, a triplet sample pair is generated to train the first image recognition model, which can improve the convergence stability and recognition accuracy of the first image recognition model. And the problem of oscillation in the training of the first image recognition model caused by randomly selecting positive sample pairs and negative sample pairs can be reduced.

[0165] Based on the same concept, this embodiment also provides a sample generation device 700, referring to Figure 7 , Figure 7 is a schematic diagram showing a sample generating device 700 according to an exemplary embodiment of the present disclosure, such as Figure 7 As shown, including:

[0166] A first determining module 701 is used to determine a plurality of image sets, each of which includes a plurality of images corresponding to the same identity information;

[0167] A first selection module 702 is used to cluster multiple images corresponding to each of the image sets to obtain multiple intra-class clusters, and select images from different intra-class clusters to form positive sample pairs, wherein the similarity between features corresponding to different images in the same intra-class cluster is less than a first similarity threshold;

[0168] A second selection module 703 is used to cluster the multiple image sets to obtain multiple inter-class clusters, and select images corresponding to different image sets in the same inter-class cluster to form negative sample pairs, wherein the similarity between the features corresponding to different image sets in the same inter-class cluster is less than a second similarity threshold;

[0169] The generating module 704 is configured to generate triplet sample pairs according to the multiple image sets, the positive sample pairs and the negative sample pairs.

[0170] Optionally, the first selection module 702 includes:

[0171] A first calculation module, used for calculating the similarity between every two images in the image set;

[0172] A first connection module, configured to regard the images in the image set as nodes, and connect the nodes corresponding to the two images whose similarity is less than the first similarity threshold, to obtain a first undirected graph;

[0173] The second determining module is used to determine the image corresponding to each subgraph in the first undirected graph as an intra-cluster of a class to obtain multiple intra-clusters.

[0174] Optionally, the first selection module 702 includes:

[0175] A first selection submodule is used to randomly select a first image from a first in-class cluster and randomly select a second image from a second in-class cluster, wherein the first in-class cluster and the second in-class cluster are different in-class clusters among the multiple in-class clusters;

[0176] The first forming module is used to form a positive sample pair with the first image and the second image.

[0177] Optionally, the second selection module 703 includes:

[0178] A second calculation module is used to calculate the similarity between the central features of every two image sets in the multiple image sets, where the central feature of the image is the image in the image set that meets a preset condition;

[0179] A second connection module, configured to regard the central features of the image sets as nodes, and connect the nodes corresponding to the central features of two image sets whose similarities are less than the second similarity threshold, to obtain a second undirected graph;

[0180] The third determination module is used to determine the image set corresponding to each subgraph in the second undirected graph as an between-class cluster to obtain the multiple between-class clusters.

[0181] Optionally, the second selection module 703 includes:

[0182] A second selection submodule is used to randomly select a third image and a fourth image from the same between-class cluster, wherein the third image and the fourth image are images corresponding to different image sets;

[0183] The second forming module is used to form a negative sample pair with the third image and the fourth image.

[0184] Optionally, the second selection module 703 includes:

[0185] A third sub-selection module is used for randomly selecting a fifth image from each image set in the same between-class cluster;

[0186] A fourth sub-selection module is used to randomly select a sixth image from each of the multiple intra-class clusters of the remaining image set, wherein the remaining image set is the image set remaining after excluding the image set in the same inter-class cluster;

[0187] The third forming module is used to form the negative sample pair with the fifth image and the sixth image.

[0188] Optionally, the first determining module is further configured to:

[0189] Determine the image set collected by the vehicle-mounted image acquisition device.

[0190] Based on the same concept, this embodiment also discloses a training device for an image recognition model, the training device comprising:

[0191] The first training module is used to train the first image recognition model according to the triplet sample pairs generated by the sample generation method disclosed in this embodiment.

[0192] Optionally, the training device further comprises:

[0193] The second training module is used to obtain a vehicle-mounted image set in a vehicle-mounted image acquisition device, and train a second image recognition model based on the vehicle-mounted image set through a normalized exponential softmax loss function to obtain the first image recognition model.

[0194] Optionally, the training device further comprises:

[0195] Image acquisition module, used to obtain RGB image sets;

[0196] An image processing module, used for performing grayscale processing on the RGB image set to obtain a grayscale image set;

[0197] The third training module is used to pre-train the third image recognition model according to the grayscale image set to obtain the second image recognition model.

[0198] Based on the same concept, this embodiment also provides a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the sample generation method and the image recognition model training method provided in this embodiment are implemented.

[0199] Based on the same concept, this embodiment also provides a controller, including:

[0200] a memory having a computer program stored thereon;

[0201] The processor is used to execute the computer program in the memory to implement the steps of the sample generation method and the image recognition model training method provided in this embodiment.

[0202] In another exemplary embodiment, a controller is also provided. The controller may be an integrated circuit (IC) or a chip, wherein the integrated circuit may be an IC or a collection of multiple ICs; the chip may include but is not limited to the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), SOC (System on Chip, SoC), etc. The above-mentioned integrated circuit or chip can be used to execute executable instructions (or codes) to implement the above-mentioned sample generation method and image recognition model training method. The executable instructions may be stored in the integrated circuit or chip, or may be obtained from other devices or equipment, for example, the integrated circuit or chip includes a processor, a memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-mentioned sample generation method and image recognition model training method are implemented; alternatively, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution, so as to implement the above-mentioned sample generation method and image recognition model training method.

[0203] Based on the same concept, this embodiment also provides a vehicle, including the controller provided by this embodiment.

[0204] Figure 8800 is a block diagram of a vehicle 800 according to an exemplary embodiment. For example, the vehicle 800 may be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 800 may be an autonomous vehicle or a semi-autonomous vehicle.

[0205] Reference Figure 8 , the vehicle 800 may include various subsystems, for example, an infotainment system 810, a perception system 820, a decision control system 830, a drive system 840, and a computing platform 850. The vehicle 800 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 800 may be interconnected by wire or wireless means.

[0206] In some embodiments, the infotainment system 810 may include a communication system, an entertainment system, and a navigation system, etc.

[0207] The perception system 820 may include several sensors for sensing information about the environment around the vehicle 800. For example, the perception system 820 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera.

[0208] The decision control system 830 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.

[0209] The drive system 840 may include components that provide powered motion for the vehicle 800. In one embodiment, the drive system 840 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting energy provided by the energy source into mechanical energy.

[0210] Some or all functions of the vehicle 800 are controlled by a computing platform 850. The computing platform 850 may include at least one processor 851 and a memory 852, and the processor 851 may execute instructions 853 stored in the memory 852.

[0211] The processor 851 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processor (Graphic Process Unit, GPU), a field programmable gate array (Field Programmable Gate Array, FPGA), a system on chip (System on Chip, SOC), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC) or a combination thereof.

[0212] The memory 852 may be implemented by any type of volatile or nonvolatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0213] In addition to instructions 853 , memory 852 may also store data, such as road maps, route information, vehicle location, direction, speed, etc. The data stored in memory 852 may be used by computing platform 850 .

[0214] In the embodiment of the present disclosure, the processor 851 can execute instruction 853 to complete all or part of the steps of the above-mentioned sample generation method and image recognition model training method.

[0215] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by a processor, the steps of the above-mentioned sample generation method and the training method of the image recognition model are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 852 including program instructions, and the above-mentioned program instructions can be executed by the processor 851 of the vehicle 800 to complete the above-mentioned sample generation method and the training method of the image recognition model. In another exemplary embodiment, a computer program product is also provided, which includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned sample generation method and the training method of the image recognition model when executed by the programmable device.

[0216] In another exemplary embodiment, a computer program product is also provided, which includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned sample generation method and image recognition model training method when executed by the programmable device.

[0217] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings; however, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, a variety of simple modifications can be made to the technical solution of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0218] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0219] In addition, various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A sample generation method, characterized in that: include: Determine a plurality of image sets, each of the image sets comprising a plurality of images corresponding to the same identity information; For each of the image sets, clustering multiple images corresponding to the image set to obtain multiple in-class clusters, randomly selecting a first image from a first in-class cluster, and randomly selecting a second image from a second in-class cluster, and forming a positive sample pair with the first image and the second image, wherein the first in-class cluster and the second in-class cluster are different in-class clusters among the multiple in-class clusters, and the similarity between features corresponding to different images in the same in-class cluster is less than a first similarity threshold; Clustering the multiple image sets to obtain multiple inter-class clusters, selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs, wherein the similarity between the features corresponding to the different image sets in the same inter-class cluster is less than a second similarity threshold; A triplet of sample pairs is generated according to the multiple image sets, the positive sample pairs, and the negative sample pairs.

2. The sample generation method according to claim 1, characterized in that: The clustering of the multiple images corresponding to the image set to obtain multiple intra-class clusters includes: Calculate the similarity between every two images in the image set; Considering the images in the image set as nodes, connecting the nodes corresponding to the two images whose similarity is less than the first similarity threshold, to obtain a first undirected graph; The image corresponding to each subgraph in the first undirected graph is determined as an intra-cluster of a class, so as to obtain a plurality of intra-cluster of the class.

3. The sample generation method according to claim 1, characterized in that: The clustering of the plurality of image sets to obtain a plurality of inter-class clusters comprises: Calculating the similarity between the central features of every two image sets in the multiple image sets, the central features of the image sets being images in the image sets that meet a preset condition; The central features of the image sets are regarded as nodes, and the nodes corresponding to the central features of the two image sets whose similarities are less than the second similarity threshold are connected to obtain a second undirected graph; The image set corresponding to each subgraph in the second undirected graph is determined as an inter-class cluster to obtain the multiple inter-class clusters.

4. The sample generation method according to claim 1, characterized in that: The selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs includes: Randomly selecting a third image and a fourth image from the same between-class cluster, wherein the third image and the fourth image are images corresponding to different image sets; The third image and the fourth image form a negative sample pair.

5. The sample generation method according to claim 1, characterized in that: The selecting images corresponding to different image sets in the same inter-class cluster to form negative sample pairs includes: For each image set in the same between-class cluster, randomly selecting a fifth image from the image set; Randomly select a sixth image from each of the multiple intra-class clusters of the remaining image set, wherein the remaining image set is the image set remaining after excluding the image set in the same inter-class cluster; The fifth image and the sixth image constitute the negative sample pair.

6. The sample generation method according to any one of claims 1 to 5, characterized in that: The determining of a plurality of image sets comprises: A plurality of image sets captured by the vehicle-mounted image capture device are determined.

7. A method for training an image recognition model, characterized in that: The method comprises: The first image recognition model is trained according to the triplet sample pairs generated by the sample generation method according to any one of claims 1 to 6.

8. The method for training an image recognition model according to claim 7, characterized in that: Before training the first image recognition model, the method further includes: A vehicle-mounted image set in a vehicle-mounted image acquisition device is obtained, and a second image recognition model is trained according to the vehicle-mounted image set by using a normalized exponential softmax loss function to obtain the first image recognition model.

9. The method for training an image recognition model according to claim 8, characterized in that: Before training the second image recognition model by using the normalized exponential softmax loss function to obtain the first image recognition model, the method further includes: Get RGB image set; Performing grayscale processing on the RGB image set to obtain a grayscale image set; The third image recognition model is pre-trained according to the grayscale image set to obtain the second image recognition model.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 9 are implemented.

11. A controller, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 9.

12. A vehicle, characterized in that: Includes the controller as claimed in claim 11.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Data processing method and device and computer equipment

    CN110163265A

  • Feature extraction model training method, data processing method, device and equipment

    CN115130536A