Image detection model training method, device, equipment, storage medium and vehicle
By generating a three-dimensional representation of the target object through the neural radiance field algorithm and training the image detection model, the problem of perspective dependence is solved and the accuracy and robustness of the image detection model are improved.
Patent Information
- Application Number
- CN202310332499.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2043-03-30
AI Technical Summary
Existing image detection models cannot accurately identify target objects under certain viewing angles, resulting in inaccurate detection results.
The preset two-dimensional image is processed through the Neural Radiance Field (NeRF) algorithm to generate a three-dimensional representation of the target object. The target two-dimensional image is determined based on multiple preset perspectives. The loss value is calculated using the detection results of the initial image detection model and actual information. The perspective corresponding to the maximum loss value is determined for training to generate the target image detection model.
The training samples of the image detection model are enriched, and the detection accuracy and robustness of the model under different perspectives are improved.
Smart Images

Figure CN116416494B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image detection, in particular to the field of automotive intelligent application technology, and specifically to a training method, apparatus, equipment, storage medium and vehicle for an image detection model. Background Art
[0002] Currently, image detection has been widely used in scenarios such as autonomous driving and surveillance. Target objects in these scenarios can usually be identified based on image detection models. However, since image detection models cannot accurately identify target objects from certain perspectives, inaccurate image detection results may result.
[0003] Related techniques primarily involve manually collecting two-dimensional images of sample objects at several fixed viewing angles and using these images as training data for image detection models, hoping to enable the image detection models to accurately identify target objects from a wider range of viewing angles. However, due to the small number of manually collected viewing angles in related techniques and their inability to accurately reflect the viewing angles that the image detection models cannot recognize, these techniques can still lead to inaccurate detection results. Summary of the Invention
[0004] This application provides a method, apparatus, device, storage medium, and vehicle for training an image detection model to at least address the technical problem of inaccurate detection results from image detection models in related technologies. The technical solution of this application is as follows:
[0005] According to the first aspect of the present application, a training method for an image detection model is provided, comprising: obtaining a preset two-dimensional image; the preset two-dimensional image includes a target object; processing the preset two-dimensional image based on a preset neural radiance fields (NeRF) algorithm to obtain a three-dimensional representation corresponding to the target object; determining multiple target two-dimensional images based on multiple preset perspectives and three-dimensional representations; a target two-dimensional image includes a target object at a preset perspective; detecting each target two-dimensional image according to an initial image detection model to obtain a detection result of the target object in each target two-dimensional image; determining a loss value of the initial image detection model for each preset perspective based on the detection result of the target object in each target two-dimensional image and actual information of the target object, to obtain multiple loss values corresponding to the multiple preset perspectives; determining the preset perspective corresponding to the maximum loss value among the multiple loss values as the target perspective, and training the initial image detection model based on the target perspective to obtain a target image detection model.
[0006] According to the above technical means, the present application can obtain a preset two-dimensional image and process the preset two-dimensional image based on the preset neural radiation field NeRF algorithm to obtain a three-dimensional representation corresponding to the target object, and then determine multiple target two-dimensional images based on multiple preset perspectives and the three-dimensional representation. Accordingly, each target two-dimensional image is detected, and the preset perspective with the largest loss value is determined from the multiple detection results, and it is determined as the target perspective for training the initial image detection model. In this way, the three-dimensional representation of the target object is reconstructed based on the neural radiation field algorithm and the preset two-dimensional image, and multiple target two-dimensional images corresponding to multiple preset perspectives are output based on the reconstructed three-dimensional representation, so that more two-dimensional images from other perspectives do not need to be manually collected, which can enrich the training samples of the image detection model. At the same time, multiple target two-dimensional images can be obtained based on multiple preset perspectives, achieving the diversity of preset perspectives. At the same time, the target perspective that maximizes the loss value of the initial image detection model is determined from multiple preset perspectives, and then the initial image detection model is trained based on the target perspective, so that the image detection model can learn the image features at the target perspective, thereby improving the accuracy of the image detection model.
[0007] In a possible implementation, the method further includes: determining a preset number of initial training rounds; determining a plurality of preset viewing angles from a preset viewing angle range according to the initial training rounds; and the number of the plurality of preset viewing angles is the number of initial training rounds.
[0008] According to the above technical means, the present application can predetermine a preset number of initial training rounds and, based on the initial training rounds, randomly generate multiple preset viewing angles from a preset viewing angle range, the same number as the initial training rounds. In this way, because the viewing angles of a model in a three-dimensional environment can be combined and varied in various ways, it is impossible to exhaustively determine the viewing angle corresponding to the two-dimensional image with the maximum loss value. Therefore, by using the preset viewing angle range and the initial training rounds, multiple preset viewing angles are determined.
[0009] In a possible embodiment, the above method further includes: updating the initial number of training rounds when multiple loss values meet preset conditions; the preset conditions include: multiple loss values are all less than a first threshold, and / or the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than a second threshold; the initial number of training rounds before updating is less than the initial number of training rounds after updating.
[0010] Based on the above technical means, the present application can determine whether multiple loss values meet preset conditions and update the initial number of training rounds if multiple loss values meet the preset conditions. In this way, by increasing the number of initial training rounds, a method for updating the initial number of training rounds is implemented, and the target perspective that maximizes the loss value of the initial image detection model can be determined, thereby improving the accuracy of the image detection model based on the target perspective.
[0011] In a possible embodiment, the above-mentioned training of the initial image detection model based on the target perspective to obtain the target image detection model includes: obtaining multiple sample two-dimensional images corresponding to the target perspective; the perspective of the object included in each sample two-dimensional image is the target perspective; using the multiple sample two-dimensional images as sample data, and using the actual information of the object included in each sample two-dimensional image as label data to train the initial image detection model to obtain the target image detection model.
[0012] According to the above technical means, the present application can obtain a plurality of sample two-dimensional images corresponding to the target perspective, use the plurality of sample two-dimensional images as sample data, and use the actual information of the object included in each sample two-dimensional image as label data to train the initial image detection model to obtain a target image detection model. In this way, by using the sample two-dimensional images under the plurality of target perspectives to train the initial image detection model to obtain the target image detection model, the target image detection model can be improved in detecting other objects under the target perspective, which is conducive to improving the robustness of the initial image detection model.
[0013] According to the second aspect provided by the present application, a training device for an image detection model is provided, including an acquisition unit, a processing unit, a determination unit, a detection unit and a training unit; the acquisition unit is used to acquire a preset two-dimensional image; the preset two-dimensional image includes a target object; the processing unit is used to process the preset two-dimensional image based on a preset neural radiation field (NeRF) algorithm to obtain a three-dimensional representation corresponding to the target object; the determination unit is used to determine multiple target two-dimensional images based on multiple preset perspectives and three-dimensional representations; a target two-dimensional image includes a target object under a preset perspective; the detection unit is used to detect each target two-dimensional image according to the initial image detection model to obtain a detection result of the target object in each target two-dimensional image; the determination unit is also used to determine the loss value of the initial image detection model for each preset perspective based on the detection result of the target object in each target two-dimensional image and the actual information of the target object, and obtain multiple loss values corresponding to multiple preset perspectives; the determination unit is also used to determine the preset perspective corresponding to the maximum loss value among the multiple loss values as the target perspective; the training unit is used to train the initial image detection model based on the target perspective to obtain a target image detection model.
[0014] According to the above technical means, the present application can obtain a preset two-dimensional image and process the preset two-dimensional image based on the preset neural radiation field NeRF algorithm to obtain a three-dimensional representation corresponding to the target object, and then determine multiple target two-dimensional images based on multiple preset perspectives and the three-dimensional representation. Accordingly, each target two-dimensional image is detected, and the preset perspective with the largest loss value is determined from the multiple detection results, and it is determined as the target perspective for training the initial image detection model. In this way, the three-dimensional representation of the target object is reconstructed based on the neural radiation field algorithm and the preset two-dimensional image, and multiple target two-dimensional images corresponding to multiple preset perspectives are output based on the reconstructed three-dimensional representation, so that more two-dimensional images from other perspectives do not need to be manually collected, which can enrich the training samples of the image detection model. At the same time, multiple target two-dimensional images can be obtained based on multiple preset perspectives, achieving the diversity of preset perspectives. At the same time, the target perspective that maximizes the loss value of the initial image detection model is determined from multiple preset perspectives, and then the initial image detection model is trained based on the target perspective, so that the image detection model can learn the image features at the target perspective, thereby improving the accuracy of the image detection model.
[0015] In a possible implementation, the determination unit is further configured to determine a preset number of initial training rounds; determine a plurality of preset viewing angles from a preset viewing angle range based on the initial training rounds; and the number of the plurality of preset viewing angles is the number of initial training rounds.
[0016] According to the above technical means, the present application can predetermine a preset number of initial training rounds and, based on the initial training rounds, randomly generate multiple preset viewing angles from a preset viewing angle range, the same number as the initial training rounds. In this way, because the viewing angles of a model in a three-dimensional environment can be combined and varied in various ways, it is impossible to exhaustively determine the viewing angle corresponding to the two-dimensional image with the maximum loss value. Therefore, by using the preset viewing angle range and the initial training rounds, multiple preset viewing angles are determined.
[0017] In one possible embodiment, the above-mentioned device also includes an updating unit; the updating unit is used to update the initial number of training rounds when multiple loss values meet preset conditions; the preset conditions include: multiple loss values are all less than a first threshold, and / or, the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than a second threshold; the initial number of training rounds before the update is less than the initial number of training rounds after the update.
[0018] Based on the above technical means, the present application can determine whether multiple loss values meet preset conditions and update the initial number of training rounds if multiple loss values meet the preset conditions. In this way, by increasing the number of initial training rounds, a method for updating the initial number of training rounds is implemented, and the target perspective that maximizes the loss value of the initial image detection model can be determined, thereby improving the accuracy of the image detection model based on the target perspective.
[0019] In one possible embodiment, the above-mentioned training unit is specifically used to: obtain multiple sample two-dimensional images corresponding to the target perspective; the perspective of the object included in each sample two-dimensional image is the target perspective; use the multiple sample two-dimensional images as sample data, and use the actual information of the object included in each sample two-dimensional image as label data to train the initial image detection model to obtain the target image detection model.
[0020] According to the above technical means, the present application can obtain a plurality of sample two-dimensional images corresponding to the target perspective, use the plurality of sample two-dimensional images as sample data, and use the actual information of the object included in each sample two-dimensional image as label data to train the initial image detection model to obtain a target image detection model. In this way, by using the sample two-dimensional images under the plurality of target perspectives to train the initial image detection model to obtain the target image detection model, the target image detection model can be improved in detecting other objects under the target perspective, which is conducive to improving the robustness of the initial image detection model.
[0021] According to the third aspect provided by the present application, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement the method of the above-mentioned first aspect and any possible implementation method thereof.
[0022] According to the fourth aspect provided by the present application, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device is enabled to execute the method in the above-mentioned first aspect and any possible implementation method thereof.
[0023] According to the fifth aspect provided by the present application, a vehicle is provided, which is equipped with a target image detection model, which is used to detect the categories of different objects, or the positions and categories; the target image detection model is trained based on the method of the above-mentioned first aspect and any possible implementation method thereof.
[0024] According to the sixth aspect provided by the present application, a computer program product is provided, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method of the above-mentioned first aspect and any possible implementation method thereof.
[0025] It should be noted that the technical effects brought about by any implementation method in the second to sixth aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here.
[0026] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application, and do not constitute an improper limitation on the present application.
[0028] Figure 1 is a flowchart of a method for training an image detection model according to an exemplary embodiment;
[0029] Figure 2 is a schematic diagram showing a method of determining a target two-dimensional image according to an exemplary embodiment;
[0030] Figure 3 is a flowchart of another method for training an image detection model according to an exemplary embodiment;
[0031] Figure 4 is a schematic diagram of another method for training an image detection model according to an exemplary embodiment;
[0032] Figure 5 is a flowchart of another method for training an image detection model according to an exemplary embodiment;
[0033] Figure 6 is a flowchart of another method for training an image detection model according to an exemplary embodiment;
[0034] Figure 7 is a block diagram of a training device for an image detection model according to an exemplary embodiment;
[0035] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0036] In order to enable ordinary people in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0037] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0038] In the following embodiments provided in this application, this application is described using electronic devices as an example.
[0039] For ease of understanding, the following detailed introduction to the training method of the image detection model provided in this application is given in conjunction with the accompanying drawings.
[0040] Figure 1 FIG. 1 is a flowchart of a method for training an image detection model according to an exemplary embodiment. Figure 1 As shown, the training method of the image detection model includes the following steps:
[0041] S101: The electronic device obtains a preset two-dimensional image.
[0042] The preset two-dimensional image includes a target object; the target object is any photographed object; the preset two-dimensional image can be a pixel-level three-primary color (red, green, blue, RGB) image, that is, the RGB image contains pixel values and grayscale values of several pixel points.
[0043] As a possible implementation manner, the electronic device obtains an image captured by an image acquisition device of the target object.
[0044] It should be noted that an image acquisition device refers to a device capable of photographing a target object. The number of preset two-dimensional images may be one or more. When there is one preset two-dimensional image, the preset two-dimensional image is obtained by the image acquisition device photographing the target object from a fixed perspective. When there are multiple preset two-dimensional images, the preset two-dimensional images are obtained by the image acquisition device photographing the target object from different perspectives.
[0045] The viewing angle refers to the angle at which the image acquisition device captures the target object.
[0046] Exemplarily, the image acquisition device may be a camera.
[0047] S102. The electronic device processes the preset two-dimensional image based on a preset neural radiation field (NeRF) algorithm to obtain a three-dimensional representation corresponding to the target object.
[0048] Among them, the neural radiation field algorithm is an algorithm for three-dimensional reconstruction through neural rendering. The main network structure of the neural radiation field is an 8-layer fully connected neural network. The input is a five-dimensional vector and the output is pixel color density and RGB color. The three-dimensional representation refers to the result obtained after three-dimensional reconstruction of the target object using the neural radiation field algorithm.
[0049] As a possible implementation method, the electronic device obtains the sparse reconstruction results, camera intrinsic parameters, camera extrinsic parameters and three-dimensional point information output for the target object through the traditional method colmap.
[0050] The electronic device samples the grid points of each preset two-dimensional image to obtain the coordinates of each pixel point, and converts the coordinates of the pixel points in the two-dimensional image to camera coordinates through the camera intrinsic parameter matrix and the camera extrinsic parameter matrix.
[0051] The electronic device unifies the light origin and direction vector in the world coordinate system, converts the camera coordinates into three-dimensional coordinates in the world coordinate system, and then preprocesses the light origin and direction vector. On each ray, dense sampling is performed near the point that ultimately contributes most to the color. The processed sample points are then mapped from low dimension to high dimension using the embedding function.
[0052] The electronic device uses the three-dimensional coordinates of each sample point in the world coordinate system and the azimuth viewing angle of each light ray corresponding to each sample point as input to the neural radiation field algorithm, and outputs the corresponding three-dimensional representation of the target object.
[0053] In actual application, the L2 loss function can be used to optimize and update the neural radiation field algorithm based on the pixel value predicted along a specific light ray and the true pixel value of the corresponding pixel point.
[0054] It should be noted that the camera intrinsic parameters include image resolution (height and width of the image) and focal length; the camera extrinsic parameters include the translation matrix and rotation matrix for converting camera coordinates to world coordinates. The translation matrix can be a 3×3 matrix, and the rotation matrix can be a 3×1 matrix; the three-dimensional point information includes the starting depth and ending depth of the light.
[0055] For example, the five-dimensional vector received by the neural radiation field is represented by (x, y, z, θ, φ), (x, y, z) represents the three-dimensional coordinates of the sample point, (θ, φ) represents the azimuth viewing angle of the sample point; the pixel color density and RGB color output by the neural radiation field are represented by (σ, c), where is the pixel color density, is an RGB color.
[0056] The neural radiation field algorithm can be instantaneous neural graph primitive (instant-NGP), MIP mapping neural radiation field (mipmapping-NeRF, Mip-NeRF), and point-based neural radiation field (Point-NeRF).
[0057] S103: The electronic device determines a plurality of target two-dimensional images based on a plurality of preset viewing angles and a three-dimensional representation.
[0058] A target two-dimensional image includes a target object under a preset viewing angle.
[0059] As a possible implementation method, the electronic device uses a neural radiation field algorithm to obtain the color density and RGB color of all pixels at each preset viewing angle, and uses the classic volume rendering method in graphics to restore the two-dimensional image at each preset viewing angle based on the color density and RGB color of all pixels, thereby obtaining multiple target two-dimensional images.
[0060] For example, Figure 2 FIG. 1 shows a schematic diagram of determining a target two-dimensional image. Figure 2 As shown in the figure, taking the target object as a camera, the preset 2D image is the image captured by the camera based on the first-person perspective. The electronic device inputs the image captured by the camera based on the first-person perspective and three preset perspectives into the neural radiance field algorithm. The neural radiance field algorithm renders the camera image at each preset perspective and outputs three target 2D images.
[0061] S104: The electronic device detects each target two-dimensional image according to the initial image detection model to obtain a detection result of the target object in each target two-dimensional image.
[0062] As a possible implementation method, the electronic device inputs each target two-dimensional image into an initial image detection model, outputs the category confidence of the target object in each target two-dimensional image, and uses the category confidence as the detection result of the target object.
[0063] In this case, the aforementioned initial image detection model is used to detect the category of the target object in the target two-dimensional image.
[0064] As another possible implementation method, the electronic device inputs each target two-dimensional image into an initial image detection model, outputs the category confidence and position box of the target object in each target two-dimensional image, and uses the category confidence and detection box as the detection result of the target object.
[0065] In this case, the aforementioned initial image detection model is used to detect the category and location of the target object in the target two-dimensional image.
[0066] It should be noted that the initial image detection model can be a single-stage image detection model or a two-stage image detection model. For example, a typical single-stage image detection model is YOLO, and a typical two-stage image detection model is Faster-RCNN.
[0067] Exemplarily, the target two-dimensional image is represented by R(V), and the detection result of the initial image detection model is represented by F(X). Then the detection result of the initial image detection model for the target two-dimensional image is F(R(V)).
[0068] S105. The electronic device calculates the loss value of the initial image detection model for each preset viewing angle based on the detection result of the target object in each target two-dimensional image and the actual information of the target object, and obtains multiple loss values corresponding to multiple preset viewing angles.
[0069] Among them, when the initial image detection model only detects the category of the target object in the target two-dimensional image, the actual information of the target object includes the true category of the target object; when the initial image detection model detects the category and position of the target object in the target two-dimensional image, the actual information of the target object includes the true category of the target object and the actual position information of the target object.
[0070] As a possible implementation method, the electronic device determines the loss value of the initial image detection model for each preset viewing angle based on the category confidence in the target object detection results in each target two-dimensional image and the true category of the target object, and obtains multiple loss values corresponding to multiple preset viewing angles.
[0071] As another possible implementation, the electronic device determines the category confidence loss value of the initial image detection model for each preset viewing angle based on the category confidence in the detection result of the target object in each target two-dimensional image and the true category of the target object, and determines the position detection loss value of the initial image detection model for each preset viewing angle based on the position box in the detection result of the target object in each target two-dimensional image and the true position information of the target object. Then, based on the category confidence loss value and position detection loss value of each preset viewing angle, the electronic device calculates the loss value of the initial image detection model for each preset viewing angle. Accordingly, multiple loss values corresponding to multiple preset viewing angles are obtained.
[0072] It should be noted that the loss value of the initial image detection model for each preset viewing angle satisfies the L(M,N) loss function, where M represents the detection result of the initial image detection model on the target two-dimensional image. The loss value of the initial image detection model for each preset viewing angle is L(F(R(V)),N).
[0073] S106. The electronic device determines a preset viewing angle corresponding to a maximum loss value among the multiple loss values as a target viewing angle.
[0074] As a possible implementation method, the electronic device sorts multiple loss values corresponding to multiple preset perspectives according to numerical size to obtain the maximum loss value, and determines the preset perspective corresponding to the maximum loss value among the multiple loss values as the target perspective for training the initial image detection model.
[0075] It should be noted that the maximum loss value among multiple loss values satisfies MAX V L(F(R(V)),N) function, where v represents the viewing angle that maximizes the loss value of the initial image detection model.
[0076] S107: The electronic device trains the initial image detection model based on the target perspective to obtain a target image detection model.
[0077] It can be understood that the technical solution provided by the embodiments of the present application obtains a preset two-dimensional image and processes the preset two-dimensional image based on a preset neural radiation field (NeRF) algorithm to obtain a three-dimensional representation corresponding to the target object. Then, based on multiple preset perspectives and the three-dimensional representation, multiple target two-dimensional images are determined. Accordingly, each target two-dimensional image is detected, and the preset perspective with the largest loss value is determined from the multiple detection results and determined as the target perspective for training the initial image detection model. In this way, a three-dimensional representation of the target object is reconstructed based on the neural radiation field algorithm and the preset two-dimensional image, and multiple target two-dimensional images corresponding to the multiple preset perspectives are output based on the reconstructed three-dimensional representation. This eliminates the need to rely on manual acquisition to obtain more two-dimensional images from other perspectives, thereby enriching the training samples of the image detection model. At the same time, multiple target two-dimensional images can be obtained based on multiple preset perspectives, achieving a diversity of preset perspectives. At the same time, the target perspective that maximizes the loss value of the initial image detection model is determined from the multiple preset perspectives. The initial image detection model is then trained based on the target perspective, enabling the image detection model to learn image features from the target perspective, thereby improving the accuracy of the image detection model.
[0078] In some embodiments, in order to determine multiple preset viewing angles, such as Figure 3 As shown, the training method of the image detection model provided in the embodiment of the present application also includes the following steps:
[0079] S201: The electronic device determines a preset number of initial training rounds.
[0080] It should be noted that the preset number of initial training rounds is set in advance in the electronic device by the operation and maintenance personnel.
[0081] S202: The electronic device determines a plurality of preset viewing angles from a preset viewing angle range according to the number of initial training rounds.
[0082] The number of preset viewing angles is the number of initial training rounds.
[0083] As a possible implementation manner, the electronic device randomly generates a plurality of preset viewing angles from a preset viewing angle range, the number of which is the same as the number of the initial training rounds.
[0084] As another possible implementation manner, the electronic device determines a plurality of preset viewing angles, which are the same in number as the initial training rounds and are evenly distributed, from a preset viewing angle range.
[0085] Exemplarily, the preset number of initial training rounds is 100, and the preset viewing angle range is 90 degrees to 160 degrees.
[0086] As will be appreciated, the technical solution provided in the embodiments of the present application predetermines a preset number of initial training rounds and, based on this initial number of training rounds, randomly generates multiple preset viewing angles from a preset viewing angle range, the same number as the initial training rounds. In this way, since the viewing angles of a model in a 3D environment can be combined and varied in various ways, it is impossible to exhaustively determine the viewing angle corresponding to the 2D image with the maximum loss value. Therefore, the preset viewing angle range and the initial number of training rounds are used to determine multiple preset viewing angles.
[0087] In some embodiments, in order to determine the target perspective that maximizes the loss value of the initial image detection model, and then improve the accuracy of the image detection model according to the target perspective, as shown in FIG. Figure 4 As shown, the training method of the image detection model provided in the embodiment of the present application also includes the following steps:
[0088] S301. The electronic device determines whether multiple loss values meet preset conditions.
[0089] The preset conditions include: the multiple loss values are all smaller than a first threshold, and / or the difference between the maximum loss value and the minimum loss value among the multiple loss values is smaller than a second threshold.
[0090] As one possible implementation, the electronic device compares each loss value with a first threshold, calculates the difference between the maximum loss value and the minimum loss value among the multiple loss values, and compares the difference with a second threshold. Accordingly, if all of the multiple loss values are less than the first threshold, and / or the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than the second threshold, the electronic device determines that the multiple loss values meet a preset condition.
[0091] In other cases, the electronic device determines that the multiple loss values do not meet the preset conditions and does not need to update the initial number of training rounds.
[0092] It can be understood that if multiple loss values are all less than the first threshold, it means that the initial image detection model can accurately detect the target object in the target two-dimensional image under multiple preset perspectives, and the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than the second threshold, it means that the overall fluctuation of the loss value of the initial image detection model for multiple preset perspectives is small, and it is necessary to re-determine the target perspective that makes the loss value of the initial image detection model the largest.
[0093] S302: When multiple loss values meet preset conditions, the electronic device updates the initial number of training rounds.
[0094] Among them, the initial number of training rounds before the update is smaller than the initial number of training rounds after the update.
[0095] In another embodiment, when multiple loss values meet a preset condition, the electronic device may further update the preset viewing angle range, for example, by increasing the preset viewing angle range or changing the preset viewing angle range.
[0096] As can be appreciated, the technical solution provided in the embodiments of the present application determines whether multiple loss values meet preset conditions and, if so, updates the initial number of training rounds. Thus, by increasing the number of initial training rounds, a method for updating the initial number of training rounds is implemented, and the target viewing angle that maximizes the loss value of the initial image detection model can be determined, thereby improving the accuracy of the image detection model based on the target viewing angle.
[0097] In some embodiments, to optimize the viewing angle, e.g. Figure 5 As shown, the training method of the image detection model provided in the embodiment of the present application also includes the following steps:
[0098] S401. Pre-training neural radiation field algorithm for electronic devices.
[0099] Among them, the pre-trained neural radiation field algorithm can refer to the implementation method of S101 to S102 above.
[0100] S402: The electronic device determines the number of training rounds, the viewing angle, and the viewing angle range, and initializes a counter.
[0101] It should be noted that the number of training rounds is represented by n, and the value of the initialization counter is 0.
[0102] S403. The electronic device determines a two-dimensional image of the target at a preset viewing angle based on a neural radiation field algorithm and a preset viewing angle.
[0103] The electronic device may determine the target two-dimensional image at the preset viewing angle based on the neural radiation field algorithm and the preset viewing angle by referring to the implementation method of S103 described above.
[0104] S404: The electronic device detects the target two-dimensional image according to the initial image detection model to obtain a detection result of the target object in the target two-dimensional image.
[0105] The electronic device detects the target two-dimensional image according to the initial image detection model, and obtains the detection result of the target object in the target two-dimensional image, which can be implemented by referring to the above S104.
[0106] S405: The electronic device calculates a loss value of the initial image detection model for a preset viewing angle based on the detection result of the target object in the target two-dimensional image.
[0107] The electronic device may calculate the loss value of the initial image detection model for each preset viewing angle based on the detection result of the target object in the target two-dimensional image, and the implementation method of the above-mentioned S105 may be referred to.
[0108] S406: The electronic device updates the optimized viewing angle according to the gradient descent of the loss function.
[0109] As a possible implementation method, the electronic device updates the preset viewing angle according to the gradient descent of the loss function.
[0110] It should be noted that updating the preset viewing angle can be achieved by changing the preset viewing angle range, or by changing the viewing angle within the preset viewing angle range.
[0111] For example, the preset viewing angle range may be increased or changed.
[0112] S407: The electronic device determines whether the number of the determined target two-dimensional images is less than a preset number of training rounds.
[0113] S408 : When the number of target two-dimensional images is less than the preset number of training rounds, the electronic device updates the counter and continues to execute the above steps S403 to S406 .
[0114] It should be noted that the update counter satisfies the following formula 1:
[0115] The value of the counter after updating = the value of the counter before updating + 1 Formula 1
[0116] S409. When the number of target two-dimensional images is greater than or equal to a preset number of training rounds, the electronic device determines a preset viewing angle corresponding to a maximum loss value among all loss values as a target viewing angle for training an initial image detection model.
[0117] In some embodiments, in order to train a target detection model, such as Figure 6 As shown, in the training method of the image detection model provided in the embodiment of the present application, the above S107 includes the following steps:
[0118] S501: The electronic device obtains a plurality of sample two-dimensional images corresponding to a target viewing angle.
[0119] The viewing angle of the object included in each sample two-dimensional image is the target viewing angle.
[0120] As a possible implementation manner, the electronic device obtains a plurality of sample two-dimensional images of different objects at the same viewing angle as the target.
[0121] For example, different objects may be vehicles, pedestrians, etc., and the target viewing angle may be 30 degrees upward.
[0122] S502: The electronic device uses a plurality of sample two-dimensional images as sample data and actual information of the object included in each sample two-dimensional image as label data to train an initial image detection model to obtain a target image detection model.
[0123] As a possible implementation method, the electronic device inputs each sample two-dimensional image into the initial image detection model, adjusts the parameters of the initial image detection model based on the detection results output by the initial image detection model and combined with the actual information of the object included in each sample two-dimensional image, to obtain the target image detection model.
[0124] It is understandable that the technical solution provided in the embodiments of the present application obtains multiple sample two-dimensional images corresponding to the target perspective, uses the multiple sample two-dimensional images as sample data, and uses the actual information of the object included in each sample two-dimensional image as label data to train the initial image detection model to obtain the target image detection model. In this way, by using the sample two-dimensional images under multiple target perspectives to train the initial image detection model to obtain the target image detection model, the target image detection model's ability to detect other objects under the target perspective can be improved, which is conducive to improving the robustness of the initial image detection model.
[0125] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to achieve the above functions, the training device or electronic device of the image detection model includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0126] The embodiment of the present application can, according to the above method, exemplarily divide the functional modules of the training device or electronic device of the image detection model. For example, the training device or electronic device of the image detection model may include various functional modules corresponding to the various functional divisions, or two or more functions may be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0127] Figure 7 FIG. 1 is a block diagram of a training device for an image detection model according to an exemplary embodiment. Figure 7 The training device 600 of the image detection model includes: an acquisition unit 601, a processing unit 602, a determination unit 603, a detection unit 604 and a training unit 606.
[0128] The acquisition unit 601 is configured to acquire a preset two-dimensional image; the preset two-dimensional image includes a target object.
[0129] The processing unit 602 is used to process the preset two-dimensional image based on the preset neural radiation field NeRF algorithm to obtain a three-dimensional representation corresponding to the target object.
[0130] The determining unit 603 is configured to determine a plurality of target two-dimensional images based on a plurality of preset viewing angles and three-dimensional representations; a target two-dimensional image includes a target object at a preset viewing angle.
[0131] The detection unit 604 is configured to detect each target two-dimensional image according to the initial image detection model to obtain a detection result of the target object in each target two-dimensional image.
[0132] The determination unit 603 is also used to determine the loss value of the initial image detection model for each preset perspective based on the detection results of the target object in each target two-dimensional image and the actual information of the target object, and obtain multiple loss values corresponding to multiple preset perspectives.
[0133] The determining unit 603 is further configured to determine a preset viewing angle corresponding to a maximum loss value among the multiple loss values as a target viewing angle.
[0134] The training unit 606 is used to train the initial image detection model based on the target perspective to obtain a target image detection model.
[0135] Optional, such as Figure 7 As shown, the determining unit 603 provided in this embodiment of the present application is further configured to:
[0136] Determine the preset number of initial training rounds.
[0137] According to the number of initial training rounds, a plurality of preset viewing angles are determined from a preset viewing angle range; the number of the plurality of preset viewing angles is the number of initial training rounds.
[0138] Optional, such as Figure 7 As shown, the training device 600 of the image detection model provided in the embodiment of the present application also includes an updating unit 605; the updating unit 605 is used to update the initial number of training rounds when multiple loss values meet preset conditions; the preset conditions include: multiple loss values are all less than a first threshold, and / or, the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than a second threshold; the initial number of training rounds before the update is less than the initial number of training rounds after the update.
[0139] Optional, such as Figure 7 As shown, the training unit 606 provided in the embodiment of the present application is specifically used to: obtain multiple sample two-dimensional images corresponding to the target perspective; the perspective of the object included in each sample two-dimensional image is the target perspective; use the multiple sample two-dimensional images as sample data, and the actual information of the object included in each sample two-dimensional image as label data to train the initial image detection model to obtain the target image detection model.
[0140] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0141] Figure 8 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 8 As shown, the electronic device 700 includes but is not limited to: a processor 701 and a memory 702 .
[0142] The memory 702 is used to store executable instructions of the processor 701. It is understandable that the processor 701 is configured to execute instructions to implement the training method of the image detection model in the above embodiment.
[0143] It should be noted that those skilled in the art can understand that Figure 8 The electronic device structure shown in the figure does not limit the electronic device, and the electronic device may include Figure 8 More or fewer components may be shown, or certain components may be combined, or the components may be arranged differently.
[0144] The processor 701 is the control center of the electronic device. It connects the various parts of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 702 and calling data stored in the memory 702, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 701 may include one or more processing units. Optionally, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly handles wireless communications. It is understood that the above-mentioned modem processor may not be integrated into the processor 701.
[0145] The memory 702 can be used to store software programs and various data. The memory 702 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and application programs required by at least one functional module (such as a determination unit, a processing unit, etc.). Furthermore, the memory 702 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0146] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 702 including instructions, and the above instructions can be executed by the processor 701 of the electronic device 700 to implement the training method of the image detection model in the above embodiment.
[0147] In actual implementation, Figure 7 The functions of the acquisition unit 601, the processing unit 602, the determination unit 603, the detection unit 604, the updating unit 605 and the training unit 606 in the embodiment can be realized by Figure 8 The processor 701 in the embodiment calls the computer program stored in the memory 702. The specific execution process can be referred to the description of the training method of the image detection model in the above embodiment, which will not be repeated here.
[0148] Optionally, the computer-readable storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0149] In an exemplary embodiment, a vehicle is also provided, which is deployed with the above-mentioned target image detection model, which is used to detect the categories of different objects, or the position and category; the target image detection model is trained based on the training method of the above-mentioned image detection model.
[0150] In an exemplary embodiment, the present application also provides a computer program product comprising one or more instructions, which can be executed by the processor 701 of the electronic device to complete the training method of the image detection model in the above embodiment.
[0151] It should be noted that when the instructions in the above-mentioned computer-readable storage medium or one or more instructions in the computer program product are executed by the processor of the electronic device, the various processes of the embodiment of the training method of the above-mentioned image detection model are implemented, and the same technical effect as the training method of the above-mentioned image detection model can be achieved. To avoid repetition, they will not be repeated here.
[0152] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete the full classification or partial functions described above.
[0153] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0154] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0155] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0156] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the full classification part or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute the full classification part or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks or optical disks.
[0157] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A training method for an image detection model, characterized in that: include: Acquire a preset two-dimensional image; the preset two-dimensional image includes a target object; Processing the preset two-dimensional image based on a preset neural radiation field (NeRF) algorithm to obtain a three-dimensional representation corresponding to the target object; Determining a plurality of target two-dimensional images based on a plurality of preset viewing angles and the three-dimensional representation; a target two-dimensional image includes the target object at a preset viewing angle; Detecting each target two-dimensional image according to the initial image detection model to obtain a detection result of the target object in each target two-dimensional image; Determining a loss value of the initial image detection model for each preset viewing angle based on a detection result of the target object in each target two-dimensional image and actual information of the target object, to obtain a plurality of loss values corresponding to the plurality of preset viewing angles; The preset viewing angle corresponding to the maximum loss value among the multiple loss values is determined as the target viewing angle, and the initial image detection model is trained based on the target viewing angle to obtain a target image detection model.
2. The method according to claim 1, characterized in that The method further comprises: Determine the preset number of initial training rounds; According to the initial training round number, the plurality of preset viewing angles are determined from a preset viewing angle range; the number of the plurality of preset viewing angles is the initial training round number.
3. The method according to claim 2, characterized in that The method further comprises: If the multiple loss values meet preset conditions, the initial number of training rounds is updated; the preset conditions include: the multiple loss values are all less than a first threshold, and / or the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than a second threshold; the initial number of training rounds before updating is less than the initial number of training rounds after updating.
4. The method according to any one of claims 1 to 3, characterized in that The initial image detection model is trained based on the target perspective to obtain a target image detection model, including: Acquire a plurality of sample two-dimensional images corresponding to the target perspective; the perspective of the object included in each sample two-dimensional image is the target perspective; The multiple sample two-dimensional images are used as sample data, and the actual information of the object included in each sample two-dimensional image is used as label data to train the initial image detection model to obtain the target image detection model.
5. A training device for an image detection model, characterized in that: include: Acquisition unit, processing unit, determination unit, detection unit and training unit; The acquisition unit is configured to acquire a preset two-dimensional image; the preset two-dimensional image includes a target object; The processing unit is configured to process the preset two-dimensional image based on a preset neural radiation field (NeRF) algorithm to obtain a three-dimensional representation corresponding to the target object; The determining unit is configured to determine a plurality of target two-dimensional images based on a plurality of preset viewing angles and the three-dimensional representation; a target two-dimensional image includes the target object at a preset viewing angle; The detection unit is configured to detect each target two-dimensional image according to the initial image detection model to obtain a detection result of the target object in each target two-dimensional image; The determining unit is further configured to determine a loss value of the initial image detection model for each preset viewing angle based on a detection result of the target object in each target two-dimensional image and actual information of the target object, thereby obtaining a plurality of loss values corresponding to the plurality of preset viewing angles; The determining unit is further configured to determine a preset viewing angle corresponding to a maximum loss value among the multiple loss values as a target viewing angle; The training unit is used to train the initial image detection model based on the target perspective to obtain a target image detection model.
6. The device according to claim 5, characterized in that The determining unit is further configured to: Determine the preset number of initial training rounds; According to the initial training round number, the plurality of preset viewing angles are determined from a preset viewing angle range; the number of the plurality of preset viewing angles is the initial training round number.
7. The device according to claim 6, characterized in that The apparatus further comprises an updating unit; The updating unit is configured to update the initial number of training rounds when the multiple loss values meet a preset condition; The preset conditions include: the multiple loss values are all less than a first threshold, and / or the difference between the maximum loss value and the minimum loss value among the multiple loss values is less than a second threshold; the initial number of training rounds before updating is less than the initial number of training rounds after updating.
8. The device according to any one of claims 5 to 7, characterized in that The training unit is specifically used for: Acquire a plurality of sample two-dimensional images corresponding to the target perspective; the perspective of the object included in each sample two-dimensional image is the target perspective; The multiple sample two-dimensional images are used as sample data, and the actual information of the object included in each sample two-dimensional image is used as label data to train the initial image detection model to obtain a target image detection model.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that When the computer-executable instructions stored in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can perform the method according to any one of claims 1 to 4.
11. A vehicle, characterized in that: The vehicle is deployed with a target image detection model, which is used to detect the categories of different objects, or the positions and categories; the target image detection model is trained based on the method described in any one of claims 1-4.
Citation Information
Patent Citations
Model training method, visual angle image generation method, device, equipment and medium
CN115409949A
Face reconstruction method, system and device based on generated image and storage medium
CN115457097A