Model training method and device, computer equipment and computer readable storage medium
By determining and using the first distribution characteristics of pixels in the image area to train the object detection model, the fitting difficulty caused by independent optimization of the parameter of the object detection model prediction box is solved, and the detection accuracy is improved.
Patent Information
- Application Number
- CN202411856187.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
AI Technical Summary
In image object detection, the prediction box parameters of the object detection model are independently optimized, resulting in difficulty in fitting, thereby reducing detection accuracy.
By acquiring an image sample, the object detection model is used to perform object detection processing, the first distribution feature of pixels in the image area indicated by the prediction border information is determined, and the object detection model is trained based on this feature.
This method avoids independent optimization of each parameter in the prediction border information, reduces the convergence difficulty of the model, and improves the detection accuracy of the trained object detection model for the target object.
Smart Images

Figure CN120014402A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of neural network technology, and in particular to a model training method, apparatus, computer device and computer-readable storage medium. Background Art
[0002] In object detection in images, the object detection model can output the parameters of the prediction box. The object detection model is trained based on the gap between the parameters of the prediction box data and the corresponding true parameters. Since each parameter is optimized independently, the loss values determined based on different parameters are different, which makes it difficult to fit the object detection model and the object detection model has poor detection accuracy for target objects. Summary of the invention
[0003] The embodiments of the present application provide a model training method, apparatus, computer device and computer-readable storage medium, which can improve the accuracy of a trained target detection model in detecting target objects.
[0004] In order to achieve the above object, according to the first aspect of the present application, a model training method is provided, comprising:
[0005] Get image samples;
[0006] Performing target detection processing on the image sample through a target detection model to determine predicted bounding box information corresponding to the target object in the image sample;
[0007] Determine a first distribution feature of pixels in the image area indicated by the predicted border information;
[0008] The target detection model is trained based on the first distribution feature to obtain a trained target detection model.
[0009] According to a second aspect of the present application, a model training device is provided, comprising:
[0010] An acquisition unit, used for acquiring image samples;
[0011] A detection unit, configured to perform target detection processing on the image sample through a target detection model to determine predicted bounding box information corresponding to the target object in the image sample;
[0012] A determining unit, configured to determine a first distribution feature of pixels in the image area indicated by the predicted border information;
[0013] A training unit is used to train the target detection model based on the first distribution feature to obtain a trained target detection model.
[0014] According to a third aspect of the present application, a computer device is provided, comprising a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any one of the model training methods provided in the embodiments of the present application.
[0015] According to a third aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store a computer program, and the computer program is loaded by a processor to execute any one of the model training methods provided in the embodiments of the present application.
[0016] The embodiment of the present application obtains image samples: performs target detection processing on the image samples through a target detection model to determine predicted bounding box information corresponding to the target object in the image samples; determines a first distribution feature of pixels in the image area indicated by the predicted bounding box information; and trains the target detection model based on the first distribution feature to obtain a trained target detection model, which can avoid independently optimizing the model for each parameter in the predicted bounding box information, reduce the difficulty of model convergence, and improve the accuracy of the trained target detection model in detecting target objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative work.
[0018] Figure 1 is a flow chart of the model training method provided in an embodiment of the present application;
[0019] Figure 2 It is a schematic diagram of a projection section provided in an embodiment of the present application;
[0020] Figure 3 is another projection section schematic diagram provided in an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of a model training device provided in an embodiment of the present application;
[0022] Figure 5 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0024] The embodiments of the present application provide a model training method, apparatus, computer equipment and computer-readable storage medium. The model training apparatus can be integrated in a computer equipment.
[0025] It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments.
[0026] This embodiment will be described from the perspective of a model training device, which can be specifically integrated into a computer device.
[0027] A model training method provided in the embodiment of the present application is as follows: Figure 1 As shown, the specific process of the model training method can be as follows:
[0028] 101. Obtain an image sample.
[0029] The image samples are images used to train the target detection model. The image samples may be plane images or curved images, such as spherical images or quasi-spherical images, that is, the image samples are spherical image samples.
[0030] 102. Perform target detection processing on the image sample through a target detection model to determine predicted bounding box information corresponding to the target object in the image sample.
[0031] Among them, the target detection model can be a neural network model used to detect target objects in images. The target detection model can include a convolutional neural network module, a self-attention module, etc. The target detection model can also be a model built based on algorithms such as YOLOv5. The structure of the target detection model can be flexibly set according to the needs of the application scenario, and is not limited here.
[0032] The target object is an object that needs to be detected from the image, and can be flexibly set according to the application scenario. For example, it can be a pedestrian, a vehicle, an obstacle, etc.
[0033] In one embodiment, the target detection model may output offset information relative to a preset anchor box, and the predicted border information is determined based on the offset information and the preset anchor box. The preset anchor box may be a predefined border. That is, in one embodiment, the step of “performing target detection processing on the image sample through the target detection model to determine the predicted border information corresponding to the target object in the image sample” includes:
[0034] Performing target detection processing on the image sample by the target detection model to obtain offset information, wherein the offset information indicates the offset degree between the predicted bounding box and the target preset anchor box;
[0035] Determine predicted bounding box information of the predicted bounding box according to the anchor box information of the target preset anchor box and the offset information.
[0036] The target detection model is used to perform target detection processing on the image sample to obtain offset information. The offset information may include the offset corresponding to each parameter in the parameters required to be included in the predefined border information to indicate the degree of offset between the predicted border and the target preset anchor frame between various parameters. For example, the border information needs to include the coordinates, length and width of the center of the projected section, then the offset information may include the offset for the coordinates, length and width of the center of the section, and the offset information corresponds to the predicted border information.
[0037] According to the anchor frame information of the target preset anchor frame and the offset information, the predicted frame information of the predicted frame is determined. For example, the anchor frame information of the target preset anchor frame is (θ α ,φ α , α α , β α , γ α ), the offset information is The predicted bounding box information (θ * ,φ * , α * , β * , γ * ).
[0038]
[0039] The offset information of the target preset anchor box and the real box (θ, φ, α, β, γ) is as follows:
[0040] t θ =(θ-θ α ) / α α , t φ =(φ-φ α ) / β α
[0041] tα =log(α / α α ), t β =log(β / β α )
[0042] t γ =γ-γ α
[0043] When training the target detection model for the image sample, the target detection model determines the target preset anchor frame according to the similarity between the preset anchor frame and the real frame, and then outputs the offset information based on the target preset anchor frame. That is, in one embodiment, before the step of "performing target detection processing on the image sample by the target detection model to obtain the offset information", the model training method provided in the embodiment of the present application may also include:
[0044] According to the anchor frame information of each preset anchor frame and the real frame information, calculating the similarity between each preset anchor frame and the real frame corresponding to the real frame information;
[0045] Determine a screening threshold corresponding to the real border information;
[0046] The target preset anchor frame is determined according to the similarity between each of the preset anchor frames and the real frame, and the screening threshold.
[0047] According to the anchor frame information of each preset anchor frame and the real frame information, the similarity is determined by cosine similarity, Euclidean distance, Manhattan distance, Spearman correlation coefficient, etc., and the similarity between each preset anchor frame and the real frame corresponding to the real frame information is calculated.
[0048] Alternatively, the corresponding third distribution feature may be determined based on the anchor box information, and the similarity between the first distribution feature and the second distribution feature may be determined based on the second distribution feature corresponding to the real border information. That is, in one embodiment, the step of "calculating the similarity between each preset anchor box and the real border corresponding to the real border information according to the anchor box information of each preset anchor box and the real border information" may include:
[0049] For each of the preset anchor frames, determining a third distribution feature of pixels in the image area indicated by the predicted border information;
[0050] The similarity between each of the preset anchor boxes and the real bounding box is determined according to the second distribution feature corresponding to the real bounding box and the third distribution feature corresponding to the preset anchor box.
[0051] For each of the preset anchor frames, the third distribution characteristics of the pixels in the image area indicated by the predicted border information are determined. The specific process can refer to the first distribution characteristics of the pixels in the image area indicated by the predicted border information in step 103, which will not be repeated here.
[0052] The similarity between the first distribution feature and the second distribution feature can be determined by cosine similarity, Euclidean distance, Manhattan distance, Spearman correlation coefficient, etc. Optionally, the similarity can also be determined by information divergence.
[0053] The similarity between the real bounding box and each preset anchor box is determined, and the target preset anchor box can be screened out according to a screening threshold. For example, the preset anchor box with a similarity greater than the screening threshold can be determined as the target preset anchor box.
[0054] The screening threshold may be predetermined, or may be determined based on the similarity between each preset anchor box and the real box. That is, in one embodiment, the step of “determining the screening threshold corresponding to the real box information” includes:
[0055] Determine the degree of similarity dispersion and the average similarity according to the similarity between each of the preset anchor frames and the real frame;
[0056] A screening threshold corresponding to the real border information is calculated according to the similarity dispersion degree and the average similarity.
[0057] According to the similarity between each preset anchor frame in multiple preset anchor frames and the real frame, the similarity dispersion degree and the average similarity of multiple similarities are determined. The average similarity can be directly calculated based on the similarity, or the average similarity can be calculated after normalizing the similarity. For example, the similarity can be determined based on the information divergence, then the average similarity It can be calculated based on the following formula, where f i,j (D kl ) is the similarity between the i-th real bounding box and the j-th preset anchor box, D kl is the information divergence between the i-th real bounding box and the j-th preset anchor box, and f() is the normalization function.
[0058]
[0059] The number of true bounding boxes is related to the number of target objects contained in the image sample.
[0060] The similarity dispersion degree may include the variance, standard deviation, etc. of multiple similarities. In one embodiment, the similarity dispersion degree It can be calculated based on the following formula.
[0061]
[0062] According to the similarity dispersion degree and the average similarity, a screening threshold corresponding to the real border information is calculated, which can be specifically determined based on the following calculation formula or by other methods, for example, the difference between the average similarity and the similarity dispersion degree can be used as the screening threshold.
[0063]
[0064] In one embodiment, the target detection model includes feature extraction
[0065] The network and decoder, the step of "performing target detection processing on the spherical image sample through the target detection model to determine the predicted bounding box information corresponding to the target object in the spherical image sample" may include:
[0066] Extracting features from the image sample through the feature extraction network to obtain image feature information of the image sample;
[0067] The decoder determines predicted bounding box information of the target object in the image sample based on the image feature information.
[0068] The feature extraction network can be a feature pyramid network (FPN), which combines the bottom-level high-resolution features with the top-level abstract semantic features by building feature pyramids at different levels, thereby improving the network's detection performance for targets of different scales. The basic structure of the FPN network can include a bottom-up feature extraction path and a top-down feature propagation path. The bottom-up path extracts features through different convolutional layers, while the top-down path passes high-level features to the lower layers through upsampling and fusion operations. There are lateral connections at each level to fuse the bottom and top features, so that the FPN network can better handle targets of different scales and sizes, thereby improving detection performance.
[0069] The decoder is used as a detection head. Its function is to generate the category and location information of the target from the image features generated by the FPN network, and map the high-level semantic features to the location and category of the target. Specifically, the softmax function can be used for classification.
[0070] 103. Determine a first distribution feature of pixels in the image area indicated by the predicted border information.
[0071] The predicted border information may be used to determine a boundary box in the image sample, that is, the predicted border information may indicate an image region in the image sample.
[0072] The first distribution feature can characterize the distribution characteristics of pixels in the image area indicated by the predicted border information. For example, it can indicate the pixel center point from which the pixels in the image area are distributed and the dispersion distance, etc. In one embodiment, the first distribution feature can be a Gaussian distribution.
[0073] In the embodiment of the present application, the image area is represented based on the first distribution feature, which improves the flexibility of representing the target position and shape, and the target detection model is trained based on the second distribution feature, and the frame information can be predicted as a whole to train the target detection model, which can reduce the sensitivity of the model to parameter fitting, thereby improving the detection accuracy. If the target detection model is trained based on the gap between the parameters in the predicted frame information and the actual parameters, the target detection model is sensitive to outliers, which will cause the target detection model to overfit these outliers and have low detection accuracy.
[0074] Optionally, the first distribution feature may be other distributions besides Gaussian distribution, such as Laplace distribution, Chi-Squared distribution, etc., which may be specifically set according to the scenario and is not limited here.
[0075] In one embodiment, a transformation matrix from a rectangular coordinate system corresponding to the image sample to a coordinate system determined based on the image area indicated by the predicted border information may be determined according to the predicted border information, and then based on the transformation matrix, a first distribution feature may be determined. That is, in one embodiment, the step of “determining the first distribution feature of pixels in the image area indicated by the predicted border information” includes:
[0076] According to the predicted border information, a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system is determined, where the target coordinate system is a coordinate system corresponding to a projection section of the image area indicated by the predicted border information:
[0077] Determine a first distribution feature of pixels in the image area indicated by the predicted border information based on the transformation matrix.
[0078] If the image sample is a plane image, the rectangular coordinate system corresponding to the image sample can be a rectangular coordinate system with the image center of the image sample as the coordinate origin, and the target coordinate system can be a rectangular coordinate system with the center of the image area indicated by the predicted border information (hereinafter referred to as the predicted image area) as the coordinate origin and the direction of the length and width of the regional image area as the coordinate axis direction.
[0079] If the image sample is a spherical image sample, the corresponding rectangular coordinate system can be a three-dimensional rectangular coordinate system with the center of the sphere of the spherical image sample as the origin, and the target coordinate system can be a rectangular coordinate system with the center of the projection section of the image area indicated by the predicted frame information as the coordinate origin and the direction of the length and width of the projection section as the coordinate axis direction. The center of the projection section of the image area indicated by the predicted frame information can be as follows Figure 2 The plane marked by the green frame shown in (a) can be considered as the plane obtained by unfolding the image area in the spherical image sample.
[0080] The coordinates of the center p of the projection section in the spherical coordinate system are P(θ, φ), and the horizontal field of view of the image area indicated by the predicted border information observed at the center of the sphere is α, and the vertical field of view is β, as shown in Figure 2 As shown in (b), the length of the projected section in the AB direction (which can be considered as the width) is ω = 2tan(0.5α), as Figure 2 As shown in (c), the length in the CD direction (which can be regarded as the length) is h=2tan(0.5β).
[0081] For Figure 3 The reference position of the projection section shown in (c) can be Figure 3 As shown in (b), Figure 3 The projection section shown in (c) is rotated to Figure 3 The projection section shown in (b) needs to be rotated by an angle γ, and the value range of γ is [-90°, 90°).
[0082] Since the rectangular coordinate system and the target coordinate system corresponding to the image sample are known, the transformation matrix of the rectangular coordinate system and the target coordinate system corresponding to the image sample can be determined. In one embodiment, the predicted frame information also includes the rotation angle of the projection section relative to the reference position, that is, γ mentioned above. The step of "determining the transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system according to the predicted frame information" may include:
[0083] Determining a first transformation matrix according to a rotation angle of the projection section relative to a reference position;
[0084] Based on the first transformation matrix, a transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system is determined.
[0085] Based on the rotation angle of the projection section relative to the reference position, the first transformation matrix is determined by the following formula:
[0086]
[0087] Based on the first transformation matrix, a transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system can be determined. For example, if the center of the section of the projection section is located at a preset position of the image sample, the first transformation matrix can be used as a transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system. That is, in one embodiment, the step of “determining the transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system based on the first transformation matrix” includes:
[0088] If the center of the projection section is located at a preset position of the image sample, the first transformation matrix is used as a transformation matrix from the rectangular coordinate system to the target coordinate system.
[0089] The preset position may be a pre-set position. For example, for a spherical image sample, the preset position may be located at the pole. The pole may refer to two points at both ends of the spherical image sample. For example, Figure 3 At the location where the center of the section shown in (a) is located, there is another pole at the other end of the spherical image sample.
[0090] If the center of the projected section is not located at the preset position of the image sample, before the step of "determining the transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system based on the first transformation matrix", the model training method provided in the embodiment of the present application may further include:
[0091] If the center of the projection section is not located at the preset position of the image sample, determining a second transformation matrix from the rectangular coordinate system to a reference coordinate system according to the first coordinate, the reference coordinate system being a coordinate system corresponding to the reference position;
[0092] The step of “determining a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system based on the first transformation matrix” includes:
[0093] A transformation matrix from the rectangular coordinate system to a target coordinate system is determined according to the first transformation matrix and the second transformation matrix.
[0094] The coordinates of the center of the projection section are (θ i ,φ i ), a second transformation matrix from the rectangular coordinate system to a reference coordinate system is determined according to the first coordinate, and the reference coordinate system may be a section plane Π i (θ i ,φ i , ω, h, 0) corresponding to the coordinate system, the second transformation matrix T can be calculated according to the following expression: i , Y i , Z i], where T represents the transposition process.
[0095]
[0096] The second transformation matrix can be expressed from Figure 3 The rectangular coordinate system corresponding to the image sample shown in (b) is Figure 3 The transformation matrix of the coordinate system corresponding to the projection section shown in (b) is, Figure 3 The rotation matrix of the projection section shown in (c) can be considered as Figure 3 (a) and Figure 3 The combination of the rotation matrices corresponding to the projection section of (b), therefore, the transformation matrix from the rectangular coordinate system to the target coordinate system can be determined according to the first transformation matrix and the second transformation matrix.
[0097] After determining the corresponding transformation matrix from the rectangular coordinate system to the target coordinate system, the first distribution feature of the pixels in the predicted image area can be determined based on the transformation matrix according to the type of the first distribution feature. In one embodiment, the predicted border information also includes the length and width of the projection section and the first coordinates of the section center of the projection section. The first distribution feature of the pixels in the image area indicated by the predicted border information is determined based on the transformation matrix, including:
[0098] Determining a covariance matrix corresponding to the predicted border information according to the length and width of the projection section;
[0099] The first distribution feature is determined according to the first coordinates of the center of the projection section, the transformation matrix and the covariance matrix.
[0100] Among them, the covariance matrix can be calculated in the following way, where ω can be considered as the length of the projection section and h can be considered as the width of the projection section.
[0101]
[0102] The first distribution feature is determined according to the first coordinate of the section center of the projection section, the transformation matrix and the covariance matrix. For example, when the first distribution feature is a Gaussian distribution, the first distribution feature based on the first coordinate of the section center, the transformation matrix and the covariance matrix can be N=(μ, ∑). If the section center of the projection section is located at a preset position, the first distribution feature can be determined according to the first coordinate of the section center of the projection section, the transformation matrix and the covariance matrix. When the first distribution feature is a Gaussian distribution, the determined first distribution feature can be N(μ0, ∑0), where μ0 can be the coordinate of the section center of the projection section in the rectangular coordinate corresponding to the image sample, ∑0=RΛR T, R is the transformation matrix, and Λ is the covariance matrix.
[0103] If the section of the projected section is not located at the preset position, the first distribution feature can be determined according to the first coordinate of the section center of the projected section, the transformation matrix and the covariance matrix. When the first distribution feature is a Gaussian distribution, the determined first distribution feature can be N(μ i ,∑ i , where μ i It can be the coordinate of the center of the projected section in the rectangular coordinates corresponding to the image sample, ∑ i =R(TΛT T )R T , the T located in the upper right corner represents the transposition processing, the T not in the upper right corner is the second transformation matrix, R is the first transformation matrix, and the determination of the first transformation matrix and the second transformation matrix can refer to the above-mentioned related content, which will not be repeated here.
[0104] In one embodiment, the image sample is a spherical image sample, and the first coordinate is the coordinate of the center of the projection section in the spherical coordinate system. The first coordinate in the spherical coordinate system can be converted into a second coordinate, and then the first distribution feature is determined in combination with the transformation matrix and the covariance matrix, that is, the step of "determining the first distribution feature according to the first coordinate of the center of the projection section, the transformation matrix and the covariance matrix" includes:
[0105] Performing coordinate conversion processing on the first coordinate to obtain a second coordinate of the center of the projection section in a rectangular coordinate system corresponding to the spherical image sample;
[0106] The first distribution feature is determined according to the second coordinates, the transformation matrix, and the covariance matrix.
[0107] Assume that the first coordinate in the section of the projection section is P(θ0, φ0, 1), perform coordinate transformation on the first coordinate to obtain the section center of the projection section. The second coordinate in the rectangular coordinate system corresponding to the spherical image sample is μ0 = [sin(φ0)cos(θ0), sin(φ0)sin(θ0), cos(φ0)].
[0108] Assume that the first coordinate of the projection section is P(θ i ,φ i , 1), coordinate transformation is performed on the first coordinate to obtain the center of the projection section, and the second coordinate in the rectangular coordinate system corresponding to the spherical image sample is μ i =(sin(φ i )cos(θ i ), sin(φ i )sin(θi ), cos(φ i )).
[0109] To determine the first distribution feature according to the second coordinates, the transformation matrix and the covariance matrix, you can refer to the above-mentioned related content and will not elaborate on it here.
[0110] 104. Train the target detection model based on the first distribution feature to obtain a trained target detection model.
[0111] Since the first distribution feature characterizes the distribution characteristics of pixels in the image area indicated by the predicted border information (hereinafter referred to as the predicted image area), it is possible to determine whether the predicted border output by the target detection model is accurate based on the first distribution feature. The target detection model is trained based on the first distribution feature, which can include self-supervised training, unsupervised training or supervised training. When the target detection model meets the preset training conditions, the trained target detection model is obtained. The preset training conditions may include the number of training times reaching a preset number threshold, or the detection accuracy of the target detection model reaching a preset accuracy threshold, etc.
[0112] In one embodiment, supervised training may be performed on the target detection model based on the first distribution feature. For example, the target detection model is trained by the difference between the first distribution feature and the second distribution feature corresponding to the real border information corresponding to the image sample. That is, the step of “training the target detection model based on the first distribution feature” includes:
[0113] Acquire a second distribution feature corresponding to real frame information corresponding to the target object in the image sample;
[0114] Calculating a loss function for the object detection model based on a similarity between the first distribution feature and the second distribution feature;
[0115] The target detection model is trained based on the loss function.
[0116] Among them, the real border information can be considered as the label of the image sample. Based on the real border information, the real image area where the target object is located can be determined. The second distribution feature can characterize the distribution characteristics of pixels in the image area indicated by the real border information. The second distribution feature can be pre-set, and its setting process can refer to the process of determining the first distribution feature.
[0117] Based on the similarity between the first distribution feature and the second distribution feature, a loss function for the target detection model is calculated; based on the loss function, the target detection model is trained until a preset training condition is met, and the preset training condition may include the number of training times reaching a preset number threshold, or the detection accuracy of the target detection model reaching a preset accuracy threshold, etc.
[0118] The similarity between the first distribution feature and the second distribution feature can be determined by cosine similarity, Euclidean distance, Manhattan distance, Spearman correlation coefficient, etc. Optionally, the similarity between the first distribution feature and the second distribution feature can also be determined by information divergence. That is, in one embodiment, before the step of "calculating the loss function for the target detection model based on the similarity between the first distribution feature and the second distribution feature", the method further includes:
[0119] calculating an information divergence between the first distribution feature and the second distribution feature;
[0120] The information divergence is normalized to obtain the similarity between the first distribution feature and the second distribution feature.
[0121] Among them, information divergence can also be called KL divergence (Kullback-Leibler divergence) or relative entropy.
[0122] For the first distribution feature being Gaussian distribution, the second distribution feature is also Gaussian distribution. Assuming that the first distribution feature is N p (μ p ,∑ p ), the second distribution feature is N g (μ g ,∑ g ), the first distribution feature N p (μ p ,∑ p ) and the second distribution feature N g (μ g ,∑ g ) can be calculated by the following formula.
[0123]
[0124] The KL divergence may be directly used as the similarity. Optionally, in order to avoid the KL divergence having a large numerical range, which may cause difficulty in model convergence, the information divergence may be normalized to obtain the similarity between the first distribution feature and the second distribution feature. Specifically, the information divergence may be normalized using the following formula.
[0125] Determine the loss function L based on information divergence reg It can be determined by the following formula, where f() represents the normalization function, which is used to normalize the information divergence D kl , τ is a hyperparameter.
[0126]
[0127] The trained target detection model can be used to perform target detection on images. For example, the trained target detection model can be deployed on the vehicle's on-board chip to perform target detection on images acquired by the vehicle through the trained target detection model to achieve autonomous driving.
[0128] In one embodiment, the vehicle may include an image acquisition module and an on-board chip, wherein the image acquisition module may obtain an image by shooting with a camera, for example, a fisheye lens may be used to shoot the environment to obtain a spherical image, which may be transmitted to the on-board chip through a communication protocol.
[0129] The on-board chip receives the image sent by the image acquisition module, processes the received image through the trained target detection model, and outputs accurate target category and predicted bounding box information.
[0130] Optionally, the trained target detection model can be applied to the autonomous driving system. The on-board chip will feed back the recognized person and obstacle information to the autonomous driving system. The autonomous driving system can effectively avoid pedestrians and obstacles by planning the path and controlling the entire vehicle system, thereby realizing the visual autonomous driving function. This fusion strategy enables the vehicle to make intelligent driving decisions based on real-time target detection results, improving the safety and efficiency of the entire autonomous driving system.
[0131] As can be seen from the above, the embodiments of the present application obtain image samples; perform target detection processing on the image samples through a target detection model to determine the predicted bounding box information corresponding to the target object in the image sample; determine the first distribution feature of pixels in the image area indicated by the predicted bounding box information; train the target detection model based on the first distribution feature to obtain a trained target detection model. By training the target model based on the first distribution feature, it is avoided that each parameter in the predicted bounding box information is independently optimized for the model, which can reduce the difficulty of convergence of the model and improve the accuracy of the trained target detection model in detecting the target object.
[0132] In order to more clearly illustrate the model training method provided by the present application, the following will take as an example an image sample that is a spherical image sample, a first distribution feature that is a Gaussian distribution feature, and border information that includes the deflection angle, length, width, and coordinates of the center of the projection section relative to the reference position.
[0133] Obtain spherical image samples, perform target detection processing on the spherical image samples through the target detection model, and determine the predicted frame information (θ * ,φ * , α * , β * , γ).
[0134] The predicted bounding box information corresponding to the target object (θ * ,φ * , α * , β * , γ) can be determined specifically through the following steps.
[0135] Determine the KL divergence between the third distribution feature corresponding to each preset anchor frame in the plurality of preset anchor frames and the second distribution feature corresponding to the real frame, determine the similarity dispersion degree and the average similarity of the plurality of KL divergences, and the average similarity Similarity dispersion By filtering threshold It can be calculated based on the following formula, where f i,j (D kl ) is the similarity between the i-th real bounding box and the j-th preset anchor box, D kl is the information divergence between the i-th real bounding box and the j-th preset anchor box, and f() is the normalization function.
[0136]
[0137] The normalization function f() is as follows, where c is a hyperparameter. The screening threshold for selecting the preset anchor box of the target is dynamically calculated based on the statistical characteristics of all normalized distances, so that the target detection model can better adapt to targets of different sizes and categories, thereby improving the detection accuracy.
[0138]
[0139] Determine the first distribution feature N based on the predicted border information p (μ p ,∑ p ), specifically, if the projection section Π indicated by the predicted border information o (θ o ,φ o , ω, h, γ) is located at the pole of the spherical image sample, then the first distribution feature is N(μ0, ∑0), where μ μ It can be the coordinate of the center of the projection section in the rectangular coordinate corresponding to the image sample, which can be determined by the following formula: ∑0 = RΛR T, R is the transformation matrix, Λ is the covariance matrix, where ω=2tan(0.5α), h=2tan(0.5β).
[0140] μx=[sin(φ0)cos(θ0), sin(φ0)sin(θ0), cos(φ0)]
[0141]
[0142] If the projection section Π indicated by the predicted bounding box information i (θ i ,φ i , ω, h, γ) is not located at the pole of the spherical image sample, then the first distribution feature is N(μ i ,∑ i ), where ∑i=R(TΛT T )R T , T=[X i , Y i , Z i ", T can be determined by the following formula, T can be based on the section Π i (θ i ,φ i ,ω,h,0)determined by the transformation matrix.
[0143] μ i =(sin(φ i )cos(θ i ), sin(φ i )sin(θ i ), coS(φ i ))
[0144]
[0145] Get the second distribution feature N corresponding to the real border information of the spherical image sample g (μ g ,∑ g ), for the first distribution feature N p (μ p ,∑ p ) and the second distribution feature N g (μ g ,∑ g ), the KL divergence of the first distribution feature and the second distribution feature can be determined by the following calculation formula.
[0146]
[0147] In order to avoid the KL divergence having a too large numerical range, which may lead to difficulty in model convergence, the information divergence may be normalized to obtain the similarity between the first distribution feature and the second distribution feature. Specifically, the information divergence may be normalized using the following formula.
[0148] Determine the loss function L based on information divergence reg It can be determined by the following formula, where f() represents the normalization function, which is used to normalize the information divergence D kl , τ is a hyperparameter.
[0149]
[0150] The target detection model is trained based on the loss function to obtain a trained target detection model.
[0151] In this embodiment of the present application, the regression loss of the target detection model is determined by Gaussian distribution, and the target preset anchor frame is determined, so that the consistency of sample selection and regression loss is achieved, which can avoid the sample scale imbalance problem caused by the inconsistency between sample selection and regression loss, and improve the performance of the target detection model.
[0152] In order to facilitate better implementation of the model training method provided in the embodiment of the present application, a model training device is also provided in one embodiment. The meanings of the terms are the same as those in the above-mentioned model training method, and the specific implementation details can refer to the description in the method embodiment.
[0153] The model training device can be integrated into a computer device, such as Figure 4 As shown, the model training device may include: an acquisition unit 301, a detection unit 302, a determination unit 303 and a training unit 304, which are specifically as follows:
[0154] (1) An acquisition unit 301 is used to acquire image samples.
[0155] (2) A detection unit 302 is used to perform target detection processing on the above image samples through a target detection model to determine predicted bounding box information corresponding to the target object in the above image samples.
[0156] In one embodiment, the detection unit 302 may also be used to:
[0157] Performing target detection processing on the image sample by the target detection model to obtain offset information, where the offset information indicates the offset degree between the predicted bounding box and the target preset anchor box;
[0158] According to the anchor frame information of the target preset anchor frame and the offset information, the predicted bounding box information of the predicted bounding box is determined.
[0159] In one embodiment, the detection unit 302 may also be used to:
[0160] According to the anchor frame information of each preset anchor frame and the real frame information, the similarity between each preset anchor frame and the real frame corresponding to the real frame information is calculated;
[0161] Determine a screening threshold corresponding to the above-mentioned real border information;
[0162] The target preset anchor frame is determined according to the similarity between each of the preset anchor frames and the real bounding box, and the screening threshold.
[0163] In one embodiment, the detection unit 302 may also be used to:
[0164] Determine the degree of similarity dispersion and the average similarity according to the similarity between each of the preset anchor boxes and the real box;
[0165] According to the above similarity dispersion degree and the above average similarity, a screening threshold corresponding to the above real border information is calculated.
[0166] In one embodiment, the detection unit 302 may also be used to:
[0167] For each of the preset anchor frames, determining a third distribution feature of pixels in the image area indicated by the predicted border information;
[0168] The similarity between each preset anchor box and the real bounding box is determined according to the second distribution feature corresponding to the real bounding box and the third distribution feature corresponding to the preset anchor box.
[0169] (3) A determination unit 303, configured to determine a first distribution feature of pixels in the image region indicated by the predicted border information.
[0170] In one embodiment, the determining unit 303 may also be used to:
[0171] Determine, according to the predicted border information, a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system, where the target coordinate system is a coordinate system corresponding to a projection section of the image area indicated by the predicted border information;
[0172] Based on the conversion matrix, a first distribution feature of pixels in the image area indicated by the predicted border information is determined.
[0173] In one embodiment, the predicted border information further includes the length and width of the projection section and the first coordinate of the section center of the projection section. The determination unit 303 may also be used to:
[0174] Determine the covariance matrix corresponding to the predicted border information according to the length and width of the projection section;
[0175] The first distribution feature is determined according to the first coordinate of the center of the projection section, the transformation matrix and the covariance matrix.
[0176] In one embodiment, the predicted border information further includes a rotation angle of the projection section relative to the reference position, and the determination unit 303 may also be used to:
[0177] Determine a first transformation matrix according to the rotation angle of the projection section relative to the reference position;
[0178] Based on the first transformation matrix, a transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system is determined.
[0179] In one embodiment, the determining unit 303 may also be used to:
[0180] If the center of the projection section is located at a preset position of the image sample, the first transformation matrix is used as a transformation matrix from the rectangular coordinate system to the target coordinate system.
[0181] In one embodiment, the determining unit 303 may also be used to:
[0182] If the center of the projection section is not located at the preset position of the image sample, a second transformation matrix from the rectangular coordinate system to a reference coordinate system is determined according to the first coordinate, and the reference coordinate system is a coordinate system corresponding to the reference position;
[0183] According to the first transformation matrix and the second transformation matrix, a transformation matrix from the rectangular coordinate system to the target coordinate system is determined.
[0184] In one embodiment, the image sample is a spherical image sample, the first coordinate is the coordinate of the center of the projection section in the spherical coordinate system, and the determination unit 303 can also be used to:
[0185] Performing coordinate conversion processing on the first coordinate to obtain the second coordinate of the center of the projection section in the rectangular coordinate system corresponding to the spherical image sample;
[0186] The first distribution feature is determined according to the second coordinate, the transformation matrix and the covariance matrix.
[0187] (4) A training unit 304, used to train the target detection model based on the first distribution feature to obtain a trained target detection model.
[0188] In one embodiment, the training unit 304 may also be used to:
[0189] Obtaining a second distribution feature corresponding to the real frame information corresponding to the target object in the above image sample;
[0190] Based on the similarity between the first distribution feature and the second distribution feature, calculating a loss function for the target detection model;
[0191] The above target detection model is trained based on the above loss function.
[0192] In one embodiment, the training unit 304 may also be used to:
[0193] Calculating the information divergence between the first distribution feature and the second distribution feature;
[0194] The information divergence is normalized to obtain the similarity between the first distribution feature and the second distribution feature.
[0195] In one embodiment, the first distribution characteristic is a Gaussian distribution.
[0196] In one embodiment, the image samples are spherical image samples.
[0197] As can be seen from the above, the embodiment of the present application acquires image samples through the acquisition unit 301; the detection unit 302 performs target detection processing on the image samples through the target detection model to determine the predicted border information corresponding to the target object in the image sample; the determination unit 303 determines the first distribution feature of the pixels in the image area indicated by the predicted border information; the training unit 304 trains the target detection model based on the first distribution feature to obtain a trained target detection model, which can improve the accuracy of the trained target detection model in detecting the target object.
[0198] The present application also provides a computer device, such as Figure 5 As shown, it shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0199] The computer device may include components such as a processor 1001 with one or more processing cores, a memory 1002 with one or more computer-readable storage media, a power supply 1003, and an input unit 1004. Those skilled in the art will appreciate that Figure 5 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0200] The processor 1001 is the control center of the computer device. It uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 1002 and calling data stored in the memory 1002, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 1001 may include one or more processing cores; preferably, the processor 1001 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and computer programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1001.
[0201] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002. The memory 1002 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, a computer program required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 1002 may also include a memory controller to provide the processor 1001 with access to the memory 1002.
[0202] The computer device also includes a power supply 1003 for supplying power to various components. Preferably, the power supply 1003 can be logically connected to the processor 1001 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 1003 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0203] The computer device may further include an input unit 1004, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0204] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 1001 in the computer device will load the executable files corresponding to the processes of one or more computer programs into the memory 1002 according to the following instructions, and the processor 1001 will run the computer programs stored in the memory 1002, thereby realizing various functions, as follows:
[0205] Get image samples;
[0206] Performing target detection processing on the image samples through the target detection model to determine the predicted bounding box information corresponding to the target object in the image samples;
[0207] Determine a first distribution feature of pixels in the image area indicated by the predicted border information;
[0208] The target detection model is trained based on the first distribution feature to obtain a trained target detection model.
[0209] The specific implementation of the above operations can be found in the previous embodiments and will not be described in detail here.
[0210] As can be seen from the above, the embodiments of the present application obtain image samples; perform target detection processing on the image samples through a target detection model to determine the predicted bounding box information corresponding to the target object in the image sample; determine the first distribution feature of the pixels in the image area indicated by the predicted bounding box information; and train the target detection model based on the first distribution feature to obtain a trained target detection model, which can improve the accuracy of the trained target detection model in detecting the target object.
[0211] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementations in the above-mentioned embodiments.
[0212] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by a computer program, or by controlling related hardware through a computer program. The computer program may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0213] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program can be loaded by a processor to execute any model training method provided in the embodiment of the present application.
[0214] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0215] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0216] Since the computer program stored in the computer-readable storage medium can execute any one of the model training methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any one of the model training methods provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0217] The above is a detailed introduction to a model training method, device, computer equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A model training method, characterized in that: include: Get image samples; Performing target detection processing on the image sample through a target detection model to determine predicted bounding box information corresponding to the target object in the image sample; Determine a first distribution feature of pixels in the image area indicated by the predicted border information; The target detection model is trained based on the first distribution feature to obtain a trained target detection model.
2. The method according to claim 1, characterized in that The determining a first distribution feature of pixels in the image area indicated by the predicted border information comprises: Determine, according to the predicted border information, a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system, wherein the target coordinate system is a coordinate system corresponding to a projection section of the image area indicated by the predicted border information; Determine a first distribution feature of pixels in the image area indicated by the predicted border information based on the transformation matrix.
3. The method according to claim 2, characterized in that The predicted border information also includes the length and width of the projection section and the first coordinate of the section center of the projection section. The determining, based on the transformation matrix, the first distribution feature of pixels in the image area indicated by the predicted border information includes: Determining a covariance matrix corresponding to the predicted border information according to the length and width of the projection section; The first distribution feature is determined according to the first coordinates of the center of the projection section, the transformation matrix and the covariance matrix.
4. The method according to claim 3, characterized in that The predicted border information also includes a rotation angle of the projection section relative to a reference position, and determining a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system according to the predicted border information includes: Determining a first transformation matrix according to a rotation angle of the projection section relative to a reference position; Based on the first transformation matrix, a transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system is determined.
5. The method according to claim 4, characterized in that The step of determining a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system based on the first transformation matrix includes: If the center of the projection section is located at a preset position of the image sample, the first transformation matrix is used as a transformation matrix from the rectangular coordinate system to the target coordinate system.
6. The method according to claim 4, characterized in that Before determining the transformation matrix from the rectangular coordinate system corresponding to the image sample to the target coordinate system based on the first transformation matrix, the method further includes: If the center of the projection section is not located at the preset position of the image sample, determining a second transformation matrix from the rectangular coordinate system to a reference coordinate system according to the first coordinate, the reference coordinate system being a coordinate system corresponding to the reference position; The step of determining a transformation matrix from a rectangular coordinate system corresponding to the image sample to a target coordinate system based on the first transformation matrix includes: A transformation matrix from the rectangular coordinate system to a target coordinate system is determined according to the first transformation matrix and the second transformation matrix.
7. The method according to claim 3, characterized in that The image sample is a spherical image sample, the first coordinate is the coordinate of the center of the projection section in the spherical coordinate system, and the determining the first distribution feature according to the first coordinate of the center of the projection section, the transformation matrix and the covariance matrix includes: Performing coordinate conversion processing on the first coordinate to obtain a second coordinate of the center of the projection section in a rectangular coordinate system corresponding to the spherical image sample; The first distribution feature is determined according to the second coordinates, the transformation matrix, and the covariance matrix.
8. The method according to claim 1, characterized in that The training of the target detection model based on the first distribution feature includes: Acquire a second distribution feature corresponding to real frame information corresponding to the target object in the image sample; Calculating a loss function for the object detection model based on a similarity between the first distribution feature and the second distribution feature; The target detection model is trained based on the loss function.
9. The method according to claim 8, characterized in that Before calculating the loss function for the target detection model based on the similarity between the first distribution feature and the second distribution feature, the method further includes: calculating an information divergence between the first distribution feature and the second distribution feature; The information divergence is normalized to obtain the similarity between the first distribution feature and the second distribution feature.
10. The method according to claim 8, characterized in that The performing target detection processing on the image sample by using the target detection model to determine the predicted bounding box information corresponding to the target object in the image sample includes: Performing target detection processing on the image sample by the target detection model to obtain offset information, wherein the offset information indicates the offset degree between the predicted bounding box and the target preset anchor box; Determine predicted bounding box information of the predicted bounding box according to the anchor box information of the target preset anchor box and the offset information.
11. The method according to claim 10, characterized in that Before performing target detection processing on the image sample by the target detection model to obtain the offset information, the method further includes: According to the anchor frame information of each preset anchor frame and the real frame information, calculating the similarity between each preset anchor frame and the real frame corresponding to the real frame information; Determine a screening threshold corresponding to the real border information; The target preset anchor frame is determined according to the similarity between each of the preset anchor frames and the real frame, and the screening threshold.
12. The method according to claim 11, characterized in that The determining of the screening threshold corresponding to the real border information includes: Determine the degree of similarity dispersion and the average similarity according to the similarity between each of the preset anchor frames and the real frame; A screening threshold corresponding to the real border information is calculated according to the similarity dispersion degree and the average similarity.
13. The method according to claim 11, characterized in that The calculating, according to the anchor frame information of each preset anchor frame and the real frame information, the similarity between each preset anchor frame and the real frame corresponding to the real frame information comprises: For each of the preset anchor frames, determining a third distribution feature of pixels in the image area indicated by the predicted border information; The similarity between each preset anchor box and the real bounding box is determined according to the second distribution feature corresponding to the real bounding box and the third distribution feature corresponding to the preset anchor box.
14. The method according to any one of claims 1 to 13, characterized in that: The first distribution characteristic is Gaussian distribution.
15. The method according to any one of claims 1-6, 8-13, characterized in that: The image samples are spherical image samples.
16. A model training device, characterized in that: include: An acquisition unit, used for acquiring image samples; A detection unit, configured to perform target detection processing on the image sample through a target detection model to determine predicted bounding box information corresponding to the target object in the image sample; A determining unit, configured to determine a first distribution feature of pixels in the image area indicated by the predicted border information; A training unit is used to train the target detection model based on the first distribution feature to obtain a trained target detection model.
17. A computer device, characterized in that It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the model training method described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, which is loaded by a processor to execute the model training method described in any one of claims 1 to 15.