Model training method and device, image recognition method and device, equipment, medium and vehicle
By adding an attribute-assisted classification branch to the MapTR model and obtaining and adjusting the multi-granularity category labeling results of the training samples, the accuracy problem of the MapTR model in identifying multi-category static elements is solved, achieving higher recognition accuracy and faster training speed.
Patent Information
- Application Number
- CN202410346064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
The MapTR model is not accurate in identifying static elements or their subcategories that are higher than 3 categories.
By adding an attribute-assisted classification branch to the MapTR model, the labeling results of the first and second categories of the training sample set are obtained. The target feature vector is obtained using the feature extraction branch and input into the prediction branch and the attribute-assisted classification branch. The loss value is calculated and the parameters are adjusted until convergence to form a trained image recognition model.
The recognition accuracy of the image recognition model for multiple categories of different granularity is improved, the model training process is simplified, the training time is shortened and the recall rate is improved.
Smart Images

Figure CN120707910A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a model training, image recognition method, device, equipment, medium and vehicle. Background Art
[0002] The MapTR model provides an effective end-to-end network structure for online vector map construction, mainly using a structural behavior to model and learn high-precision maps. In existing solutions, when performing image recognition on images around the vehicle to determine static elements in the image such as lane lines, road boundaries, zebra crossings, etc., the collected image is input into the MapTR model, and image recognition is performed based on the MapTR model and the image recognition results are output, where the image recognition results include category information and point set information of the target object in the image.
[0003] However, the MapTR model was trained based on the Nuscenes dataset, which only contains three categories of static elements: zebra crossings, lane markings, and road boundaries. When identifying static elements that far exceed these three categories, or the subcategories corresponding to each static element, the MapTR model has inaccurate recognition issues. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure provides a model training, image recognition method, device, equipment, medium and vehicle.
[0005] A first aspect of the present disclosure provides a model training method, the method comprising:
[0006] Obtaining a training sample set and a labeling result corresponding to the training sample set, wherein the labeling result includes first category, second category, and point set information corresponding to the target object in the training sample, and the division granularity of the first category is greater than the division granularity of the second category;
[0007] Input the training sample set into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch, and an attribute auxiliary classification branch, and obtains the target feature vectors corresponding to the training sample set based on the feature extraction branch;
[0008] Input the target feature vector into the prediction branch and the attribute auxiliary classification branch respectively, and obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch. The first recognition result includes the first predicted category and the target point set information, and the second recognition result includes the second predicted category. The division granularity of the first predicted category is greater than or equal to the division granularity of the second predicted category.
[0009] Calculating a loss value based on the first recognition result, the second recognition result, the labeling result, and a preset loss function to obtain a target loss value of a preset image recognition model;
[0010] The parameters of the preset image recognition model are adjusted according to the target loss value until the preset loss function converges to obtain a trained image recognition model.
[0011] A second aspect of the present disclosure provides an image recognition method, the method comprising:
[0012] Obtain the image to be detected;
[0013] The image to be detected is input into the trained image recognition model, and the image to be detected is recognized based on the trained image recognition model to obtain the recognition result of at least one target object in the image to be detected. The recognition result of each target object includes the category and point set information corresponding to the target object. The trained image recognition model is obtained based on the training method described in the first aspect above.
[0014] A third aspect of the present disclosure provides a model training device, the device comprising:
[0015] A training set acquisition module is used to obtain a training sample set and the annotation results corresponding to the training sample set. The annotation results include the first category, the second category, and the point set information corresponding to the target object in the training sample. The first category has a larger division granularity than the second category.
[0016] A feature extraction module is used to input the training sample set into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch, and an attribute auxiliary classification branch, and obtains the target feature vectors corresponding to each training sample set based on the feature extraction branch;
[0017] A result prediction module is used to input the target feature vector into the prediction branch and the attribute auxiliary classification branch respectively, to obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch, wherein the first recognition result includes a first prediction category and target point set information, and the second recognition result includes a second prediction category, and the division granularity of the first prediction category is greater than or equal to the division granularity of the second prediction category;
[0018] A loss value calculation module is used to calculate the loss value based on the first recognition result, the second recognition result, the labeling result and a preset loss function to obtain a target loss value of the preset image recognition model;
[0019] The parameter adjustment module is used to adjust the parameters of the preset image recognition model according to the target loss value until the preset loss function converges to obtain a trained image recognition model.
[0020] A fourth aspect of the present disclosure provides an image recognition device, the device comprising:
[0021] An image acquisition module, used to acquire an image to be detected;
[0022] An image recognition module is used to input the image to be detected into a trained image recognition model, recognize the image to be detected based on the trained image recognition model, and obtain the recognition result of at least one target object in the image to be detected. The recognition result of each target object includes the category and point set information corresponding to the target object. The trained image recognition model is obtained based on the training method described in the first aspect above.
[0023] A fifth aspect of the present disclosure provides an electronic device, the device including:
[0024] Memory;
[0025] processor; and
[0026] A computer program, wherein the computer program is stored in a memory and is configured to be executed by a processor to implement the model training method of the first aspect or the image recognition method of the second aspect as described above.
[0027] The sixth aspect of the embodiments of the present disclosure provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the model training method of the first aspect or the image recognition method of the second aspect described above is implemented.
[0028] A seventh aspect of an embodiment of the present disclosure provides a vehicle, comprising the electronic device of the fifth aspect described above.
[0029] The technical solution provided by the embodiments of the present disclosure has the following advantages over the prior art:
[0030] The model training, image recognition method, apparatus, device, medium and vehicle provided by the embodiments of the present disclosure can obtain a training sample set and a labeling result corresponding to the training sample set, wherein the labeling result includes the first category, the second category and the point set information corresponding to the target object in the training sample, the division granularity of the first category is greater than the division granularity of the second category, the training sample set is input into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch and an attribute auxiliary classification branch, the target feature vectors corresponding to the training sample set are obtained based on the feature extraction branch, the target feature vectors are respectively input into the prediction branch and the attribute auxiliary classification branch, and a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch are obtained, the first recognition result includes a first prediction category and target point set information, The second recognition result includes a second prediction category, and the division granularity of the first prediction category is greater than or equal to the division granularity of the second prediction category. The loss value is calculated based on the first recognition result, the second recognition result, the labeling result and the preset loss function to obtain the target loss value of the preset image recognition model. The parameters of the preset image recognition model are adjusted according to the target loss value until the preset loss function converges to obtain a trained image recognition model. Therefore, during the training process of the preset image recognition model, the labeling results can include categories of different division granularities corresponding to the target objects in the training samples, and the preset image recognition model is trained based on categories of different division granularities, so that the trained image recognition model can recognize categories of multiple different division granularities, thereby improving the recognition accuracy of the trained image recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0032] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0033] Figure 1 is a flowchart of a model training method provided by an embodiment of the present disclosure;
[0034] Figure 2 is a structural diagram of an image recognition model provided by an embodiment of the present disclosure;
[0035] Figure 3 is a flowchart of an image recognition method provided by an embodiment of the present disclosure;
[0036] Figure 4 is a structural diagram of a model training device provided by an embodiment of the present disclosure;
[0037] Figure 5 is a structural diagram of an image recognition device provided by an embodiment of the present disclosure;
[0038] Figure 6 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.
[0040] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0041] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0042] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0043] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0044] Typically, existing image recognition models, such as the MapTR model, are trained based on the Nuscenes dataset, which contains only three categories of static elements: zebra crossings, lane markings, and road boundaries. This leads to inaccurate recognition of static elements exceeding these three categories, or of subcategories corresponding to each static element. To address this issue, the present disclosure provides a model training method, which is described below in conjunction with specific examples.
[0045] Figure 1 It is a flowchart of a model training method provided by an embodiment of the present disclosure. The method can be executed by a model training device. The model training device can be implemented in software and / or hardware. The model training device can be configured in an electronic device, such as a server or terminal or a server cluster. The terminal can specifically include a computer or tablet computer, a vehicle-mounted terminal, or any device that can be used to process the model training method.
[0046] like Figure 1 As shown, the model training method provided by the embodiment of the present disclosure includes the following steps.
[0047] S110 , obtaining a training sample set and a labeling result corresponding to the training sample set, wherein the labeling result includes first category, second category, and point set information corresponding to the target object in the training sample, and the division granularity of the first category is greater than the division granularity of the second category.
[0048] In an embodiment of the present disclosure, after receiving a model training instruction, the electronic device obtains a training sample set and a labeling result corresponding to the training sample set.
[0049] In the embodiment of the present disclosure, the training sample set may be a training sample set for static elements in a map, wherein the static elements may be static elements such as lane lines, road boundary lines, zebra crossings, etc.
[0050] The first category can be understood as the category obtained by coarse classification of static elements, such as lane lines, road boundaries, stop lines, zebra crossings, impassable areas, lane center lines, up and down dividing lines, lane interruption lines, intersection boundaries, etc.
[0051] The second category can be understood as a category obtained by subdividing the first category. For example, the second category corresponding to the lane line may include a single white line, a single yellow line, a double white line, a double yellow line, a left dashed line and a right solid line, etc.
[0052] The first category and the second category can be divided according to user needs and actual conditions, and are not limited here.
[0053] For example, the first category may include 11 major categories, and the second category may include 54 minor categories obtained by subdividing the 11 major categories corresponding to the first category.
[0054] In some embodiments of the present disclosure, the model training instruction includes a training sample set and the labeling results corresponding to the training sample set. After receiving the model training instruction, the electronic device parses the model training instruction to obtain the training sample set and the labeling results corresponding to the training sample set.
[0055] In other embodiments of the present disclosure, the model training instruction includes identification information corresponding to the training sample set. After receiving the model training instruction, the electronic device searches for the training sample set corresponding to the identification information from the target database based on the identification information corresponding to the training sample set, and then obtains the training sample set and the labeling results corresponding to the training sample set.
[0056] S120. Input the training sample set into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch, and an attribute auxiliary classification branch, and obtains target feature vectors corresponding to the training sample sets based on the feature extraction branch.
[0057] In an embodiment of the present disclosure, after the electronic device obtains the training sample set, it inputs the training sample set into a preset image recognition model, and obtains target feature vectors corresponding to the training sample set based on the feature extraction branch in the preset image recognition model.
[0058] In the disclosed embodiments, the preset image recognition model can be a model obtained by adding an attribute-assisted classification branch to an existing image recognition model, such as the MapTR model. The preset image recognition model can include a feature extraction branch, a prediction branch, and an attribute-assisted classification branch, and can also include an input layer and a backbone network.
[0059] In the embodiment of the present disclosure, the attribute-assisted classification branch can be understood as a branch used to assist existing image recognition models such as the MapTR model in detecting and identifying target object categories.
[0060] The attribute-assisted classification branch may include multiple attribute branches, each attribute branch is used to detect and identify the category corresponding to the branch, and each attribute branch may be composed of a convolution layer, a regularization layer, and a linear classification layer.
[0061] In the disclosed embodiments, the MapTR model may include an input layer, a backbone network, a feature extraction branch, and a prediction branch, wherein the feature extraction branch may include an encoder and a decoder. For example, the feature extraction branch may be a transformer model. The role of each network structure in MapTR is the same as that of each network result in the existing MapTR model, and will not be elaborated here.
[0062] The target feature vector is a feature vector obtained by decoding the feature vector of the bird's-eye view corresponding to the training sample.
[0063] Specifically, the specific implementation method of obtaining target feature vectors corresponding to the training sample sets based on the feature extraction branch in the preset image recognition model is similar to the implementation method of obtaining feature vectors using the existing MapTR model, and will not be repeated here.
[0064] Figure 2 is a structural diagram of an image recognition model provided by an embodiment of the present disclosure, such as Figure 2 As shown, the preset image recognition model includes a MapTR model and an attribute-assisted classification branch. The MapTR model includes an input layer, a backbone network, an encoder, a decoder, a point regression classification and a classification branch, wherein the point regression branch and the classification branch belong to the prediction branch; the attribute-assisted classification branch includes a lane line attribute branch, a lane centerline attribute branch, a road boundary attribute branch, an impassable area attribute branch, etc. The specific branches contained in the attribute-assisted classification branch are set according to the specific situation and are not limited here.
[0065] Among them, the lane line attribute branch, lane centerline attribute branch, road boundary attribute branch, and impassable area attribute branch are respectively used to detect and identify their corresponding subcategories (subcategories). For example, the lane line attribute branch detects and identifies the single white line, single yellow line, etc. corresponding to the lane line.
[0066] S130. Input the target feature vector into the prediction branch and the attribute auxiliary classification branch respectively to obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch. The first recognition result includes a first prediction category and target point set information, and the second recognition result includes a second prediction category. The division granularity of the first prediction category is greater than or equal to the division granularity of the second prediction category.
[0067] In an embodiment of the present disclosure, after obtaining the target feature vector, the electronic device inputs the target feature vector into the prediction branch and the attribute auxiliary classification branch respectively, identifies the target feature vector based on the prediction branch and outputs a first recognition result, and at the same time, identifies the target feature vector based on the attribute auxiliary classification branch and outputs a second recognition result.
[0068] The prediction branch includes a point regression branch and a classification branch, wherein the point regression branch is used to identify the target feature vector and generate target point set information corresponding to the target feature vector, and the classification branch is used to identify the target feature vector to obtain a first prediction category.
[0069] The target point set information can be understood as the point set information corresponding to the target object such as a lane line.
[0070] In some embodiments of the present disclosure, taking the existing image recognition model as the MapTR model as an example, the division granularity of the first prediction category is greater than the division granularity of the second prediction category. It can be understood that the category of the target object output by the classification branch in the prediction branch in the MapTR model is the first category, and the category of the target object output by the attribute-assisted classification branch is the second category. For example, the category of the target object output by the classification branch in the prediction branch in the MapTR model is the lane line, and the category of the target object output by the attribute-assisted classification branch is the subcategory corresponding to the lane line, such as a single white line.
[0071] In other embodiments of the present disclosure, taking the existing image recognition model as the MapTR model as an example, the division granularity of the first prediction category is equal to the division granularity of the second prediction category. It can be understood that the category of the target object output by the classification branch in the prediction branch in the MapTR model is the second category, and the category of the target object output by the attribute-assisted classification branch is also the second category. For example, the category of the target object output by the classification branch in the prediction branch in the MapTR model is the single white line corresponding to the lane line, and the category of the target object output by the attribute-assisted classification branch is also the subcategory single white line corresponding to the lane line.
[0072] S140: Calculate the loss value based on the first recognition result, the second recognition result, the labeling result, and a preset loss function to obtain a target loss value of the preset image recognition model.
[0073] In an embodiment of the present disclosure, after obtaining the first recognition result and the second recognition result, the electronic device calculates the loss value based on the first recognition result, the second recognition result, the labeling result and the preset loss function to obtain the target loss value of the preset image recognition model.
[0074] Specifically, after obtaining the first recognition result and the second recognition result, the electronic device inputs the first recognition result, the second recognition result and the labeling result into the corresponding preset loss function respectively, calculates the loss value, and then obtains the target loss value corresponding to the preset image recognition model based on the obtained loss value.
[0075] S150 , adjusting parameters of a preset image recognition model according to a target loss value until a preset loss function converges, thereby obtaining a trained image recognition model.
[0076] In an embodiment of the present disclosure, after obtaining the target loss value, the electronic device adjusts the parameters of the preset image recognition model according to the target loss value until the preset loss function converges to obtain a trained image recognition model.
[0077] In some embodiments of the present disclosure, when the division granularity of the first prediction category is greater than the division granularity of the second prediction category, the parameters of the preset image recognition model are adjusted based on the target loss value until the preset loss function converges to obtain a trained first image recognition model.
[0078] In other embodiments of the present disclosure, when the division granularity of the first prediction category is equal to the division granularity of the second prediction category, the parameters of the preset image recognition model are adjusted based on the target loss value until the preset loss function converges to obtain a trained second image recognition model.
[0079] It should be noted that the specific implementation method of adjusting the parameters of the preset image recognition model according to the target loss value is similar to the existing implementation method of adjusting the parameters of the model based on the loss value during the model training process, and will not be repeated here.
[0080] In an embodiment of the present disclosure, a training sample set and a labeling result corresponding to the training sample set can be obtained, the labeling result includes the first category, the second category and the point set information corresponding to the target object in the training sample, the division granularity of the first category is greater than the division granularity of the second category, the training sample set is input into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch and an attribute auxiliary classification branch, and target feature vectors corresponding to the training sample set are obtained based on the feature extraction branch, and the target feature vectors are respectively input into the prediction branch and the attribute auxiliary classification branch to obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch, the first recognition result includes a first prediction category and target point set information, and the second recognition result includes a second prediction category and target point set information. Category, the division granularity of the first prediction category is greater than or equal to the division granularity of the second prediction category, and the loss value is calculated according to the first recognition result, the second recognition result, the labeling result and the preset loss function to obtain the target loss value of the preset image recognition model, and the parameters of the preset image recognition model are adjusted according to the target loss value until the preset loss function converges to obtain a trained image recognition model. Therefore, in the process of training the preset image recognition model, the labeling results can include categories of different division granularities corresponding to the target objects in the training samples, and the preset image recognition model is trained based on categories of different division granularities, so that the trained image recognition model can recognize categories of multiple different division granularities, thereby improving the recognition accuracy of the trained image recognition model.
[0081] In an embodiment of the present disclosure, when the division granularity of the first prediction category is equal to the division granularity of the second prediction category, a loss value is calculated based on the first recognition result, the second recognition result, the labeling result and a preset loss function to obtain a target loss value of a preset image recognition model, which may specifically include: substituting the first prediction category in the first recognition result and the second category in the labeling result into the preset first loss function for calculation to obtain a first loss value; substituting the target point set information in the first recognition result and the point set information in the labeling result into the preset second loss function for calculation to obtain a second loss value; substituting the second prediction category in the second recognition result and the second category in the labeling result into the preset first loss function for calculation to obtain a third loss value; and fusing the first loss value, the second loss value and the third loss value to obtain a target loss value.
[0082] In some embodiments of the present disclosure, the first loss value, the second loss value, and the third loss value are fused to obtain a target loss value, which may specifically include adding and averaging the first loss value, the second loss value, and the third loss value to obtain a first average value, and determining the first average value as the target loss value.
[0083] In some embodiments of the present disclosure, the first loss value, the second loss value, and the third loss value are fused to obtain a target loss value, which may specifically include determining the weights corresponding to the first loss value, the second loss value, and the third loss value respectively according to preset weights, multiplying the first loss value, the second loss value, and the third loss value respectively by their corresponding weights, adding and averaging the obtained products to obtain a second average value, and determining the second average value as the target loss value.
[0084] In an embodiment of the present disclosure, an attribute-assisted classification branch can be added to a preset image recognition model, and the attribute-assisted classification branch can be used to assist an existing image recognition model, such as a MapTR model, in model training. The preset image recognition model parameters are adjusted based on the loss value of the attribute-assisted classification branch and the loss value of each branch in an existing image recognition model, such as a MapTR model, to obtain a trained image recognition model, so that the trained image recognition model can recognize multiple categories, while improving the accuracy of the recognized categories.
[0085] In an embodiment of the present disclosure, the attribute-assisted classification branch includes at least one attribute branch, the at least one attribute branch is determined based on the first category in the labeling result, and the second predicted category includes the predicted category output by each attribute branch.
[0086] Substituting the second predicted category and the second category in the second recognition result into the preset first loss function for calculation to obtain the third loss value can specifically include: for each training sample, determining the first category corresponding to the training sample according to the labeling result corresponding to the training sample, and determining the target attribute branch corresponding to the first category; substituting the target predicted category and the second category output by the target attribute branch into the preset first loss function for calculation to obtain the third loss value.
[0087] For example, for a certain training sample, the second category in the corresponding annotation result is the single white line, a subcategory corresponding to the lane line. At this time, the target attribute branch is determined to be the lane line attribute branch, and then the target prediction category output by the lane line attribute branch and the second category in the annotation result are substituted into the preset first loss function for calculation, and the obtained loss value is determined as the third loss value.
[0088] In the embodiment of the present disclosure, the target attribute branch corresponding to the second category can be determined according to the second category in the annotation result corresponding to each training sample, and the third loss value can be determined according to the target prediction category of the target attribute branch. The third loss value is not calculated for other attribute branches, thereby improving the accuracy of the obtained loss value.
[0089] In an embodiment of the present disclosure, when the division granularity of the first prediction category is greater than the division granularity of the second prediction category, a loss value is calculated based on the first recognition result, the second recognition result, the labeling result and a preset loss function to obtain a target loss value of a preset image recognition model, which may specifically include: substituting the first prediction category in the first recognition result and the first category in the labeling result into the preset first loss function for calculation to obtain a fourth loss value; substituting the target point set information in the first recognition result and the point set information in the labeling result into the preset second loss function for calculation to obtain a second loss value; substituting the second prediction category in the second recognition result and the second category in the labeling result into the preset first loss function for calculation to obtain a third loss value; and fusing the fourth loss value, the second loss value and the third loss value to obtain a target loss value.
[0090] In the embodiment of the present disclosure, the fourth loss value, the second loss value and the third loss value are fused to obtain the specific implementation of the target loss value, which is similar to the implementation of the first loss value, the second loss value and the third loss value to obtain the target loss value in the above embodiment of the present disclosure, and will not be repeated here.
[0091] In an embodiment of the present disclosure, the second predicted category in the second recognition result and the second category in the labeling result are substituted into the preset first loss function for calculation to obtain a third loss value, which can specifically include: substituting the predicted category and the second category corresponding to each attribute branch in the auxiliary attribute classification branch into the preset first loss function for calculation, adding and averaging the loss values corresponding to each attribute branch to obtain a third average value, and determining the third average value as the third loss value.
[0092] In the embodiment of the present disclosure, the classification branch in the prediction branch of the preset image recognition model can be trained for the first category, and the attribute auxiliary classification branch can be trained for the second category, so as to output the first category and the second category of the target object in different branches respectively. Through the hierarchical classification training method, the recognition difficulty of the preset image recognition model for a large number of categories is simplified, the convergence speed of the preset image recognition model is improved, the model training time is shortened, and at the same time, the prediction or recognition difficulty of each classification branch is reduced to a certain extent, and the recall rate is improved.
[0093] Figure 3 This is a flowchart of an image recognition method provided by an embodiment of the present disclosure. The method can be executed by an image recognition device, which can be implemented in software and / or hardware. The image recognition device can be configured in an electronic device, such as a server or terminal or a server cluster, wherein the terminal can specifically include a computer or tablet computer, a vehicle-mounted terminal, or any device that can be used to process the image recognition method.
[0094] like Figure 3 As shown, the image recognition method provided by the embodiment of the present disclosure includes the following steps:
[0095] S310: Acquire an image to be detected.
[0096] In an embodiment of the present disclosure, after receiving an image recognition request, the electronic device obtains an image to be detected based on the image recognition request.
[0097] The image to be detected may be an image of the surrounding area of the vehicle including the road.
[0098] In some embodiments of the present disclosure, after receiving an image recognition request, the electronic device controls the image acquisition device to acquire an image, thereby obtaining the image to be detected.
[0099] In other embodiments of the present disclosure, after receiving an image recognition request, the electronic device identifies the image recognition request, determines image identification information corresponding to the image recognition request, and obtains the image to be detected from the target database based on the image identification information.
[0100] S320: Input the image to be detected into the trained image recognition model, recognize the image to be detected based on the trained image recognition model, and obtain a recognition result of at least one target object in the image to be detected, wherein the recognition result of each target object includes the category and point set information corresponding to the target object.
[0101] In an embodiment of the present disclosure, the trained image recognition model may be a trained first image recognition model or a trained second image recognition model obtained based on the model training method described in the above embodiment of the present disclosure.
[0102] Specifically, after obtaining the image to be detected, the electronic device directly inputs the image to be detected into a trained image recognition model, and the trained image recognition model recognizes the image to be detected to obtain a recognition result of at least one target object in the image to be detected, wherein the recognition result includes the category and point set information corresponding to the target object.
[0103] The categories of the target object include a first category and a second category. The division granularity of the first category is greater than the division granularity of the second category. For example, the category of the target object is the subcategory single white line in the lane line, where the first category is the lane line and the second category is the single white line.
[0104] In an embodiment of the present disclosure, an image to be detected can be obtained, the image to be detected can be input into a trained image recognition model, and the image to be detected can be recognized based on the trained image recognition model to obtain a recognition result of at least one target object in the image to be detected. The recognition result of each target object includes the category and point set information corresponding to the target object, wherein the trained image recognition model is trained based on categories of different division granularities corresponding to each target object in the image, thereby improving the accuracy of the recognition results obtained based on the trained image recognition model.
[0105] Furthermore, in some embodiments of the present disclosure, when the granularity of the category division output by the prediction branch in the trained image recognition model is greater than the granularity of the category division output by the attribute auxiliary classification branch in the trained image recognition model, the image to be detected is recognized based on the trained image recognition model to obtain the recognition result of at least one target object in the image to be detected, which may specifically include: identifying the category of the target object based on the prediction branch to obtain a first category corresponding to the target object; identifying the category of the target object based on the attribute auxiliary classification branch in the trained image recognition model to obtain a second category corresponding to the target object; and determining the category corresponding to the target object based on the first category and the second category.
[0106] In an embodiment of the present disclosure, the granularity of the categories output by the prediction branch in the trained image recognition model is greater than the granularity of the categories output by the attribute auxiliary classification branch in the trained image recognition model, that is, the trained first image recognition model in the above embodiment of the present disclosure.
[0107] Determining the category corresponding to the target object based on the first category and the second category can specifically include: matching the first category with the output result of each attribute branch in the attribute auxiliary classification branch, determining the target category corresponding to the target attribute branch belonging to the first category, determining the target category as the second category, and determining the first category and the second category as the categories corresponding to the target object.
[0108] In the embodiment of the present disclosure, when the granularity of the category division output by the prediction branch in the trained image recognition model is greater than the granularity of the category division output by the attribute auxiliary classification branch in the trained image recognition model, the category corresponding to the target object is determined based on the output result of the prediction branch and the output result of the attribute auxiliary classification branch, thereby reducing the difficulty of image recognition by the trained image recognition model.
[0109] In other embodiments of the present disclosure, when the granularity of the categories output by the prediction branch in the trained image recognition model is equal to the granularity of the categories output by the attribute auxiliary classification branch in the trained image recognition model, the image to be detected is recognized based on the trained image recognition model to obtain the recognition result of at least one target object in the image to be detected, which may specifically include: recognizing the category of the target object based on the prediction branch to obtain the category corresponding to the target object.
[0110] In an embodiment of the present disclosure, the granularity of the categories output by the prediction branch in the trained image recognition model is equal to the granularity of the categories output by the attribute auxiliary classification branch in the trained image recognition model, that is, the trained second image recognition model in the above embodiment of the present disclosure. At this time, the output result of the prediction branch is directly determined as the category corresponding to the target object.
[0111] In the embodiment of the present disclosure, when the granularity of the categories output by the prediction branch in the trained image recognition model is equal to the granularity of the categories output by the attribute auxiliary classification branch in the trained image recognition model, the output result of the prediction branch in the trained image recognition model is directly determined as the category corresponding to the target object, thereby improving the recognition efficiency of the image recognition result.
[0112] Figure 4The figure is a schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure. The model training device in the embodiment of the present disclosure can be provided in an electronic device, which can be a server, a terminal, or a server cluster. The terminal can specifically include a computer, a tablet computer, an in-vehicle terminal, or any other device capable of processing the model training method, without limitation herein.
[0113] like Figure 4 As shown, the model training device 400 may include a training set acquisition module 410, a feature extraction module 420, a result prediction module 430, a loss value calculation module 440 and a parameter adjustment module 450.
[0114] The training set acquisition module 410 can be used to obtain a training sample set and the labeling results corresponding to the training sample set. The labeling results include the first category, second category and point set information corresponding to the target object in the training sample. The division granularity of the first category is greater than the division granularity of the second category.
[0115] The feature extraction module 420 can be used to input the training sample set into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch and an attribute auxiliary classification branch, and the target feature vectors corresponding to the training sample set are obtained based on the feature extraction branch.
[0116] The result prediction module 430 can be used to input the target feature vector into the prediction branch and the attribute auxiliary classification branch respectively, to obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch, wherein the first recognition result includes a first prediction category and target point set information, and the second recognition result includes a second prediction category, and the division granularity of the first prediction category is greater than or equal to the division granularity of the second prediction category.
[0117] The loss value calculation module 440 can be used to calculate the loss value based on the first recognition result, the second recognition result, the labeling result and the preset loss function to obtain the target loss value of the preset image recognition model.
[0118] The parameter adjustment module 450 can be used to adjust the parameters of the preset image recognition model according to the target loss value until the preset loss function converges to obtain a trained image recognition model.
[0119] In an embodiment of the present disclosure, a training sample set and a labeling result corresponding to the training sample set can be obtained, the labeling result includes the first category, the second category and the point set information corresponding to the target object in the training sample, the division granularity of the first category is greater than the division granularity of the second category, the training sample set is input into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch and an attribute auxiliary classification branch, and target feature vectors corresponding to the training sample set are obtained based on the feature extraction branch, and the target feature vectors are respectively input into the prediction branch and the attribute auxiliary classification branch to obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch, the first recognition result includes a first prediction category and target point set information, and the second recognition result includes a second prediction category and target point set information. Category, the division granularity of the first prediction category is greater than or equal to the division granularity of the second prediction category, and the loss value is calculated according to the first recognition result, the second recognition result, the labeling result and the preset loss function to obtain the target loss value of the preset image recognition model, and the parameters of the preset image recognition model are adjusted according to the target loss value until the preset loss function converges to obtain a trained image recognition model. Therefore, in the process of training the preset image recognition model, the labeling results can include categories of different division granularities corresponding to the target objects in the training samples, and the preset image recognition model is trained based on categories of different division granularities, so that the trained image recognition model can recognize categories of multiple different division granularities, thereby improving the recognition accuracy of the trained image recognition model.
[0120] In some embodiments of the present disclosure, the loss value calculation module 440 may include a first calculation unit, a second calculation unit, a third calculation unit, and a first determination unit.
[0121] The first calculation unit can be used to substitute the first predicted category in the first recognition result and the second category in the labeling result into a preset first loss function for calculation to obtain a first loss value.
[0122] The second calculation unit can be used to substitute the target point set information in the first recognition result and the point set information in the labeling result into a preset second loss function for calculation to obtain a second loss value.
[0123] The third calculation unit can be used to substitute the second predicted category in the second recognition result and the second category in the labeling result into the preset first loss function for calculation to obtain a third loss value.
[0124] The first determination unit can be used to fuse the first loss value, the second loss value and the third loss value to obtain a target loss value.
[0125] In some embodiments of the present disclosure, the attribute-assisted classification branch includes at least one attribute branch, and the at least one attribute branch is determined based on the first category in the labeling result.
[0126] The third calculation unit can be specifically used to determine, for each training sample, the first category corresponding to the training sample according to the labeling result corresponding to the training sample, and determine the target attribute branch corresponding to the first category; substitute the target prediction category and the second category output by the target attribute branch into the preset first loss function for calculation to obtain a third loss value.
[0127] In some embodiments of the present disclosure, the loss value calculation module 440 may further include a fourth calculation unit and a second determination unit.
[0128] The fourth calculation unit can be used to substitute the first predicted category in the first recognition result and the first category in the labeling result into the preset first loss function for calculation to obtain a fourth loss value.
[0129] The second determining unit may be configured to perform a fusion process on the fourth loss value, the second loss value, and the third loss value to obtain a target loss value.
[0130] It should be noted that Figure 4 The model training device 400 shown can execute each step in the above method embodiment and realize each process and effect in the above model training method embodiment, which will not be described in detail here.
[0131] Figure 5 This is a structural diagram of an image recognition device provided in an embodiment of the present disclosure. The image recognition device in the embodiment of the present disclosure can be set in an electronic device, and the electronic device can be a server or a terminal or a server cluster. The terminal can specifically include a computer or a tablet computer, a vehicle-mounted terminal, or any device that can be used to process the image recognition method, etc., and there is no limitation here.
[0132] like Figure 5 As shown, the image recognition device 500 may include an image acquisition module 510 and an image recognition module 520 .
[0133] The image acquisition module 510 can be used to acquire an image to be detected.
[0134] The image recognition module 520 can be used to input the image to be detected into a trained image recognition model, recognize the image to be detected based on the trained image recognition model, and obtain the recognition result of at least one target object in the image to be detected. The recognition result of each target object includes the category and point set information corresponding to the target object. The trained image recognition model is obtained based on the model training method described in the above embodiment of the present disclosure.
[0135] In an embodiment of the present disclosure, an image to be detected can be obtained, the image to be detected can be input into a trained image recognition model, and the image to be detected can be recognized based on the trained image recognition model to obtain a recognition result of at least one target object in the image to be detected. The recognition result of each target object includes the category and point set information corresponding to the target object, wherein the trained image recognition model is trained based on categories of different division granularities corresponding to each target object in the image, thereby improving the accuracy of the recognition results obtained based on the trained image recognition model.
[0136] In some embodiments of the present disclosure, the image recognition module 520 can be specifically used to identify the category of the target object based on the prediction branch to obtain a first category corresponding to the target object when the category division granularity output by the prediction branch in the trained image recognition model is greater than the category division granularity output by the attribute auxiliary classification branch in the trained image recognition model; identify the category of the target object based on the attribute auxiliary classification branch in the trained image recognition model to obtain a second category corresponding to the target object; and determine the category corresponding to the target object based on the first category and the second category.
[0137] It should be noted that Figure 5 The image recognition device 500 shown can execute the various steps in the above method embodiment and realize the various processes and effects in the above image recognition method embodiment, which will not be described in detail here.
[0138] Figure 6 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown.
[0139] In the embodiments of the present disclosure, Figure 6 The electronic device shown may be a server or a terminal or a server cluster, wherein the terminal may specifically include a computer or a tablet computer, a vehicle-mounted terminal, etc., which is not limited here.
[0140] like Figure 6 As shown, the electronic device may include a processor 610 and a memory 620 storing computer program instructions.
[0141] Specifically, the processor 610 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0142] Memory 620 may include a large-capacity memory for information or instructions. By way of example, and not limitation, memory 620 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 620 may include removable or non-removable (or fixed) media. Where appropriate, memory 620 may be internal or external to the integrated gateway device. In certain embodiments, memory 620 is non-volatile solid-state memory. In certain embodiments, memory 620 includes read-only memory (ROM). Where appropriate, the ROM may be mask-programmed ROM, programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0143] The processor 610 reads and executes the computer program instructions stored in the memory 620 to perform the steps of the model training method or image recognition method provided in the embodiments of the present disclosure.
[0144] In one example, the electronic device may further include a transceiver 630 and a bus 640. Figure 6 As shown, the processor 610 , the memory 620 and the transceiver 630 are connected via a bus 640 and communicate with each other.
[0145] The bus 640 includes hardware, software, or both. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 640 may include one or more buses.
[0146] The embodiments of the present disclosure also provide a computer-readable storage medium, which can store a computer program. When the computer program is executed by a processor, the processor implements the model training method and image recognition method provided by the embodiments of the present disclosure.
[0147] The above-mentioned storage medium may, for example, include a memory 620 of computer program instructions, and the above-mentioned instructions may be executed by the processor 610 of the electronic device to complete the model training method provided in the embodiment of the present disclosure. Optionally, the storage medium may be a non-transitory computer-readable storage medium, for example, a non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device.
[0148] The embodiments of the present disclosure further provide a vehicle, which includes electronic equipment and can implement the various processes and effects in the above embodiments of the present disclosure, which will not be described in detail here.
[0149] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a list of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or apparatus.
[0150] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. A model training method, characterized in that: The method comprises: Obtaining a training sample set and a labeling result corresponding to the training sample set, wherein the labeling result includes a first category, a second category, and point set information corresponding to a target object in the training sample, wherein a division granularity of the first category is greater than a division granularity of the second category; Inputting the training sample set into a preset image recognition model, wherein the preset image recognition model includes a feature extraction branch, a prediction branch, and an attribute auxiliary classification branch, and obtaining target feature vectors corresponding to the training sample sets based on the feature extraction branch; Inputting the target feature vector into the prediction branch and the attribute-assisted classification branch respectively, obtaining a first recognition result output by the prediction branch and a second recognition result output by the attribute-assisted classification branch, wherein the first recognition result includes a first predicted category and target point set information, and the second recognition result includes a second predicted category, and the division granularity of the first predicted category is greater than or equal to the division granularity of the second predicted category; Calculating a loss value based on the first recognition result, the second recognition result, the labeling result, and a preset loss function to obtain a target loss value of the preset image recognition model; The parameters of the preset image recognition model are adjusted according to the target loss value until the preset loss function converges to obtain a trained image recognition model.
2. The method according to claim 1, characterized in that The calculating the loss value according to the first recognition result, the second recognition result, the labeling result, and a preset loss function to obtain a target loss value of the preset image recognition model includes: Substituting the first predicted category in the first recognition result and the second category in the labeling result into a preset first loss function for calculation to obtain a first loss value; Substituting the target point set information in the first recognition result and the point set information in the labeling result into a preset second loss function for calculation to obtain a second loss value; Substituting the second predicted category in the second recognition result and the second category in the labeling result into the preset first loss function for calculation to obtain a third loss value; The first loss value, the second loss value, and the third loss value are fused to obtain the target loss value.
3. The method according to claim 2, characterized in that The attribute auxiliary classification branch includes at least one attribute branch, the at least one attribute branch is determined based on the first category in the labeling result, and the second predicted category includes the predicted category output by each attribute branch; Substituting the second predicted category and the second category in the second recognition result into the preset first loss function for calculation to obtain a third loss value includes: For each training sample, determine a first category corresponding to the training sample according to the labeling result corresponding to the training sample, and determine a target attribute branch corresponding to the first category; The target prediction category output by the target attribute branch and the second category are substituted into the preset first loss function for calculation to obtain the third loss value.
4. The method according to claim 1, wherein The calculating the loss value according to the first recognition result, the second recognition result, the labeling result, and a preset loss function to obtain a target loss value of the preset image recognition model includes: Substituting the first predicted category in the first recognition result and the first category in the labeling result into a preset first loss function for calculation to obtain a fourth loss value; Substituting the target point set information in the first recognition result and the point set information in the labeling result into a preset second loss function for calculation to obtain a second loss value; Substituting the second predicted category in the second recognition result and the second category in the labeling result into the preset first loss function for calculation to obtain a third loss value; The fourth loss value, the second loss value, and the third loss value are fused to obtain the target loss value.
5. An image recognition method, characterized in that: The method comprises: Obtain the image to be detected; The image to be detected is input into a trained image recognition model, and the image to be detected is recognized based on the trained image recognition model to obtain a recognition result of at least one target object in the image to be detected, wherein the recognition result of each target object includes the category and point set information corresponding to the target object, and the trained image recognition model is obtained based on the training method described in any one of claims 1 to 4 above.
6. The method according to claim 5, characterized in that When the granularity of the categories output by the prediction branch in the trained image recognition model is greater than the granularity of the categories output by the attribute auxiliary classification branch in the trained image recognition model, the method of recognizing the image to be detected based on the trained image recognition model to obtain a recognition result of at least one target object in the image to be detected includes: Identifying the category of the target object based on the prediction branch to obtain a first category corresponding to the target object; Identifying the category of the target object based on the attribute auxiliary classification branch in the trained image recognition model to obtain a second category corresponding to the target object; A category corresponding to the target object is determined based on the first category and the second category.
7. A model training device, characterized in that: include: A training set acquisition module is used to obtain a training sample set and a labeling result corresponding to the training sample set, wherein the labeling result includes a first category, a second category, and point set information corresponding to the target object in the training sample, and the division granularity of the first category is greater than the division granularity of the second category; a feature extraction module, configured to input the training sample set into a preset image recognition model, wherein the preset image recognition model is a model comprising a feature extraction branch, a prediction branch, and an attribute auxiliary classification branch, and to obtain target feature vectors corresponding to the training sample sets based on the feature extraction branch; A result prediction module is configured to input the target feature vector into the prediction branch and the attribute auxiliary classification branch, respectively, to obtain a first recognition result output by the prediction branch and a second recognition result output by the attribute auxiliary classification branch, wherein the first recognition result includes a first predicted category and target point set information, the second recognition result includes a second predicted category, and the division granularity of the first predicted category is greater than or equal to the division granularity of the second predicted category; a loss value calculation module, configured to calculate a loss value based on the first recognition result, the second recognition result, the labeling result, and a preset loss function to obtain a target loss value of the preset image recognition model; A parameter adjustment module is used to adjust the parameters of the preset image recognition model according to the target loss value until the preset loss function converges to obtain a trained image recognition model.
8. An image recognition device, characterized in that: include: An image acquisition module, used to acquire an image to be detected; An image recognition module is used to input the image to be detected into a trained image recognition model, identify the image to be detected based on the trained image recognition model, and obtain a recognition result of at least one target object in the image to be detected, wherein the recognition result of each target object includes the category and point set information corresponding to the target object, and the trained image recognition model is obtained based on the training method described in any one of claims 1 to 4 above.
9. An electronic device, characterized in that: include: Memory; processor; as well as computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 4 or claims 5 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 or claims 5 to 6 is implemented.
11. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 9.