Model training, target detection method, device and computer equipment
By introducing multiple second coordinate systems in three-dimensional target detection and determining the occlusion ratio of the sample detection frame at different angles, the problem of inaccurate occlusion degree labeling in the existing technology is solved, and more accurate target detection and occlusion ratio output are achieved.
Patent Information
- Application Number
- CN202511031145.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-25
AI Technical Summary
In the existing technology, three-dimensional target detection models lack an effective occlusion degree labeling method when dealing with occluded objects, resulting in insufficient detection accuracy and robustness.
Multiple second coordinate systems are introduced. Based on the mapping relationship between the first coordinate system and the second coordinate system, the occlusion ratio of the sample detection frame at different angles is determined. The initial model is trained using the sample driving scene graph and the occlusion ratio to construct a target detection model.
The accuracy of occlusion ratio annotation is improved, the detection accuracy and robustness of the target detection model are enhanced, and the detection box and its occlusion ratio can be output.
Smart Images

Figure CN120526259B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a model training, target detection method, device and computer equipment. Background Art
[0002] In the fields of computer vision and autonomous driving, 3D object detection is a key technology used to identify and locate objects in three-dimensional space. This key technology is primarily implemented through neural networks, and training data is the foundation for training neural networks. Neural networks build their models by learning the patterns and relationships between labeled data in the training data. Therefore, without high-quality training data, neural networks cannot learn effective features.
[0003] Currently, there are two main methods for acquiring training data: pure manual labeling and semi-automatic labeling using a large model for pre-brush and manual correction. For occluded objects in the real world, both manual and semi-automatic labeling methods use simple definitions of occlusion properties, typically classified as unoccluded, partially occluded, or severely occluded. This simple definition relies entirely on the subjective judgment of the annotator and is quite arbitrary. This arbitrary and abstract nature makes it difficult to add and quantify the occlusion level attribute to neural network models, which in turn affects the accuracy and robustness of detection. Summary of the Invention
[0004] Based on this, it is necessary to provide a model training, target detection method, device and computer equipment to address the above technical problems, which can construct training data containing sample occlusion ratio labels, and thus ensure that the trained target detection model can accurately predict the occlusion ratio of the target object.
[0005] In a first aspect, the present application provides a model training method, comprising:
[0006] Inputting a sample driving scene graph of a sample vehicle into an initial model, obtaining first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene graph in a first coordinate system; the first coordinate system is a vehicle coordinate system with the forward direction of the sample vehicle as the vertical axis;
[0007] Based on a mapping relationship between the first coordinate system and different second coordinate systems, second coordinate information of different sample detection frames in each second coordinate system is determined according to the first coordinate information of different sample detection frames in the first coordinate system; each second coordinate system and the first coordinate system have the same coordinate origin, and the vertical axis of each second coordinate system has a different angle with the vertical axis of the first coordinate system;
[0008] Determining the sample occlusion ratios corresponding to the different sample detection frames according to the second coordinate information of the different sample detection frames in the respective second coordinate systems;
[0009] The sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames are used to train the initial model to obtain the target detection model.
[0010] In one embodiment, the second coordinate information of each sample detection frame in each second coordinate system includes the position coordinates of the vertices of the sample detection frame in each second coordinate system;
[0011] Determining the sample occlusion ratios corresponding to the different sample detection frames according to the second coordinate information of the different sample detection frames in the respective second coordinate systems includes:
[0012] Determining candidate occlusion ratios of different sample detection frames in each second coordinate system according to the position coordinates of the vertices of different sample detection frames in each second coordinate system;
[0013] For each sample detection frame, the maximum value of the candidate occlusion ratios of the sample detection frame in each second coordinate system is used as the sample occlusion ratio corresponding to the sample detection frame.
[0014] In one embodiment, determining candidate occlusion ratios of different sample detection frames in each second coordinate system according to position coordinates of vertices of different sample detection frames in each second coordinate system includes:
[0015] For each second coordinate system, based on the position coordinates of the vertices of different sample detection frames in the second coordinate system, filter out from each sample detection frame a first detection frame whose vertex vertical coordinates are greater than zero and a second detection frame whose vertex vertical coordinates are less than zero;
[0016] Using the lower limit of the occlusion ratio as the candidate occlusion ratio of the second detection frame in the second coordinate system; and
[0017] Determine candidate occlusion ratios of the first detection frame in the second coordinate system according to position coordinates of vertices of the first detection frame in the second coordinate system.
[0018] In one embodiment, when there are at least two first detection frames, determining candidate occlusion ratios of the first detection frames in the second coordinate system according to position coordinates of vertices of the first detection frames in the second coordinate system includes:
[0019] Determine the vertical coordinate of the center point of each first detection frame according to the position coordinates of the vertices of each first detection frame in the second coordinate system;
[0020] Sort the first detection frames in descending order according to the vertical coordinates of the center points of the first detection frames to obtain serial numbers of the first detection frames;
[0021] According to the serial number of each first detection frame and the position coordinates of the vertex in the second coordinate system, the candidate occlusion ratio of each first detection frame in the second coordinate system is determined in sequence.
[0022] In one embodiment, determining the candidate occlusion ratio of each first detection frame in the second coordinate system in sequence according to the serial number of each first detection frame and the position coordinates of the vertex in the second coordinate system includes:
[0023] Projecting each first detection frame onto a vertical plane of the second coordinate system based on the position coordinates of the vertices of each first detection frame in the second coordinate system to obtain a projection surface of each first detection frame on the vertical plane; wherein the vertical plane of the second coordinate system is a plane formed by the horizontal axis and the vertical axis of the second coordinate system;
[0024] For each first detection frame, select a detection frame with a sequence number greater than the sequence number of the first detection frame from the first detection frames to serve as a reference detection frame for the first detection frame;
[0025] Determining an intersection area between a projection surface of the first detection frame and a projection surface of the reference detection frame;
[0026] Determine a candidate occlusion ratio of the first detection frame in the second coordinate system according to an intersection area of a projection surface of the first detection frame and a projection surface of the reference detection frame, and a projection area of the projection surface of the first detection frame.
[0027] In one embodiment, different second coordinate systems are constructed by:
[0028] In the bird's-eye view plane where the first coordinate system is located, a preset number of rays are emitted in sequence in a counterclockwise direction with the coordinate origin of the first coordinate system as the starting point;
[0029] A preset number of rays are used as the vertical axis of each second coordinate system, and the horizontal axes of different second coordinate systems are determined according to the vertical axis of each second coordinate system to construct each second coordinate system.
[0030] In a second aspect, the present application also provides a target detection method, comprising:
[0031] Obtain the current driving scene graph of the target vehicle;
[0032] Input the current driving scene image into the target detection model to obtain the initial detection frames output by the target detection model for the current driving scene image and the target occlusion ratio corresponding to each initial detection frame;
[0033] The initial detection frame whose target occlusion ratio in each initial detection frame is less than the preset occlusion threshold is used as the target detection frame;
[0034] According to the target detection frame, the target detection result of the target vehicle is determined.
[0035] In a third aspect, the present application also provides a model training device, comprising:
[0036] A first detection module is configured to input a sample driving scene graph of a sample vehicle into an initial model, and obtain first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene graph in a first coordinate system; the first coordinate system is a vehicle coordinate system with the sample vehicle's forward direction as its vertical axis;
[0037] a coordinate mapping module for determining, based on a mapping relationship between the first coordinate system and different second coordinate systems, second coordinate information of different sample detection frames in each second coordinate system according to the first coordinate information of different sample detection frames in the first coordinate system; each second coordinate system and the first coordinate system have the same coordinate origin, and the vertical axes of different second coordinate systems have different angles with the vertical axis of the first coordinate system;
[0038] an occlusion determination module, configured to determine the sample occlusion ratios corresponding to different sample detection frames based on the second coordinate information of the different sample detection frames in each second coordinate system;
[0039] The model training module is used to train the initial model using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames to obtain the target detection model.
[0040] In a fourth aspect, the present application further provides a target detection device, comprising:
[0041] A data acquisition module is used to obtain the current driving scene map of the target vehicle;
[0042] The occlusion prediction module is used to input the current driving scene image into the target detection model to obtain the initial detection frames output by the target detection model for the current driving scene image and the target occlusion ratio corresponding to each initial detection frame;
[0043] A data screening module is used to select the initial detection frames whose target occlusion ratio in each initial detection frame is less than a preset occlusion threshold as the target detection frame;
[0044] The second detection module is used to determine the target detection result of the target vehicle according to the target detection frame.
[0045] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps involved in the first aspect when executing the computer program.
[0046] In a sixth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps involved in the second aspect when executing the computer program.
[0047] In a seventh aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps involved in the first aspect when executed by a processor.
[0048] In an eighth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps involved in the second aspect when executed by a processor.
[0049] In a ninth aspect, the present application further provides a computer program product, comprising a computer program, which implements the steps involved in the first aspect when executed by a processor.
[0050] In a tenth aspect, the present application further provides a computer program product, comprising a computer program, which implements the steps involved in the second aspect when executed by a processor.
[0051] The above-described model training, object detection method, apparatus, and computer device introduce different second coordinate systems, each of which has the same origin as the first coordinate system and has different angles between its longitudinal axis and the longitudinal axis of the first coordinate system. This means that the different second coordinate systems divide the bird's-eye view plane of the first coordinate system. Based on the mapping relationship between the first coordinate system and the different second coordinate systems, the second coordinate information of different sample detection frames in each second coordinate system can be accurately determined based on the first coordinate information of different sample detection frames in the first coordinate system. Furthermore, the sample occlusion ratios corresponding to different sample detection frames are determined based on the second coordinate information of different sample detection frames in each second coordinate system. This is equivalent to comprehensively considering the occlusion ratios of the same sample detection frame at different angles. Compared to manually annotating occlusion ratios, this method avoids the subjectivity and arbitrariness of manual annotation and improves the accuracy of the sample occlusion ratios. Furthermore, the object detection model obtained by training the initial model using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames can output not only the detection frames but also the occlusion ratios of the detection frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 This is an application environment diagram of a model training method in one embodiment;
[0054] Figure 2 Schematic diagram of a flow chart of a model training method in one embodiment;
[0055] Figure 3 A schematic diagram of a process for determining sample occlusion ratios of different sample detection frames in one embodiment;
[0056] Figure 4 Schematic diagram of a process for determining a candidate occlusion ratio in one embodiment;
[0057] Figure 5 is a schematic diagram of a flow chart for determining a candidate occlusion ratio in another embodiment;
[0058] Figure 6A Schematic diagram of a flow chart for determining a candidate occlusion ratio in yet another embodiment;
[0059] Figure 6B Schematic diagram of the intersection of the projection plane of the first detection frame and the projection plane of the reference detection frame in one embodiment;
[0060] Figure 7 1 is a flow chart of a target detection method according to an embodiment;
[0061] Figure 8 is a structural block diagram of a model training device in one embodiment;
[0062] Figure 9 is a structural block diagram of a target detection device in one embodiment;
[0063] Figure 10 is a diagram of the internal structure of a computer device in one embodiment;
[0064] Figure 11 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0066] The model training method provided in the embodiment of the present application can be executed by a computer device with powerful computing and storage functions, for example, it can be executed by an onboard terminal in a vehicle, or it can be executed by a server. In an optional embodiment, the model training method provided in the present application can be applied to Figure 1In the application environment shown. The sample vehicle 101 is equipped with a collection device for collecting the driving scene image of the sample vehicle, which can be a sensor or a camera. The sample vehicle 101 can communicate with the server 103 via a network. The server 103 is integrated with a data storage system that can store data that the server 103 needs to process. Optionally, the server 103 inputs the sample driving scene image collected by the collection device in the sample vehicle 101 into the initial model, obtains the first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene image in the first coordinate system, and based on the mapping relationship between the first coordinate system and different second coordinate systems, determines the second coordinate information of different sample detection frames in each second coordinate system according to the first coordinate information of different sample detection frames in the first coordinate system; further, the server 103 determines the sample occlusion ratio corresponding to different sample detection frames according to the second coordinate information of different sample detection frames in each second coordinate system, and uses the sample driving scene image and the sample occlusion ratio corresponding to different sample detection frames to train the initial model to obtain a target detection model.
[0067] For example, the sample vehicle and the target vehicle (i.e., the vehicle for which object detection is required) can be the same vehicle or different vehicles. If the sample vehicle and the target vehicle are different vehicles, server 103 can transmit the object detection model to the vehicle-mounted terminal 102 in the target vehicle 104, so that the vehicle-mounted terminal 102 integrates the object detection model and performs object detection on the driving scene image captured by the vehicle's sensors. Server 103 can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0068] In an exemplary embodiment, Figure 2 As shown, a model training method is provided, which is applied to Figure 1 The server 103 in the example is used as an example to illustrate the method, which specifically includes the following steps:
[0069] S201 , inputting a sample driving scene graph of a sample vehicle into an initial model, and obtaining first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene graph in a first coordinate system.
[0070] The sample vehicle may be a vehicle equipped with a collection device. The sample driving scene image is an image of the sample vehicle's driving environment collected by the collection device of the sample vehicle. The initial model is a neural network model that only has a target detection function, and the sample detection frame is a detection frame output by the initial model for locating the target object in the sample driving scene image. In an embodiment of the present application, the sample detection frame is a three-dimensional frame, represented by a set of coordinate parameters, which is used to mark the specific position of the target object in the sample driving scene image. In this embodiment, the coordinate parameters are the first coordinate information in the first coordinate system, and the first coordinate system is a vehicle coordinate system with the sample vehicle's forward direction as the vertical axis (x-axis). The vehicle coordinate system is a coordinate system with the sample vehicle's center of mass as the coordinate origin, the sample vehicle's forward direction as the positive direction of the vertical axis, the left side of the sample vehicle as the positive direction of the horizontal axis, and the vertical direction as the positive direction of the vertical axis.
[0071] Optionally, while driving, the sample vehicle may collect driving scene images under different driving scenarios and transmit them to the server as sample driving scene images. The server inputs the sample driving scene images of the sample vehicle into the initial model. The initial model performs object detection on the sample driving scene images based on preset model parameters and outputs the first coordinate information of each sample detection box in the sample driving scene images in the first coordinate system.
[0072] S202 , based on a mapping relationship between the first coordinate system and different second coordinate systems, and according to first coordinate information of different sample detection frames in the first coordinate system, determine second coordinate information of different sample detection frames in each second coordinate system.
[0073] Each second coordinate system and the first coordinate system have the same coordinate origin, and the angles between the longitudinal axes of different second coordinate systems and the longitudinal axis of the first coordinate system are different.
[0074] In an optional embodiment, the method for constructing different second coordinate systems can be to evenly divide the bird's-eye view plane where the first coordinate system is located into a preset number of planes, and determine a preset direction in each plane as the vertical axis of the second coordinate system, and use the coordinate origin of the first coordinate system as the coordinate origin of the second coordinate system. Finally, use the three-dimensional coordinate system construction rules to determine the horizontal axis of each second coordinate system to construct different second coordinate systems.
[0075] In another optional embodiment, the method for constructing different second coordinate systems can be, in the bird's-eye view plane where the first coordinate system is located, starting from the coordinate origin of the first coordinate system, emitting a preset number of rays in a counterclockwise direction in sequence; using the preset number of rays as the longitudinal axis of each second coordinate system, and determining the transverse axis of different second coordinate systems based on the longitudinal axis of each second coordinate system to construct each second coordinate system. Exemplarily, in the bird's-eye view plane of the vehicle coordinate system, starting from the negative direction of the transverse axis (y-axis), N rays are emitted in sequence in a counterclockwise direction from the coordinate origin of the vehicle coordinate system. N is a hyperparameter that can be flexibly set according to needs. Finally, the entire bird's-eye view plane is divided into equal parts. and use each ray as the vertical axis of a second coordinate system.
[0076] The mapping relationship between the first coordinate system and any second coordinate system represents the coordinate conversion rule between the first coordinate system and the second coordinate system. In the embodiment of the present application, the mapping relationship between the first coordinate system and any second coordinate system can be expressed by the following formula (1):
[0077] (1)
[0078] in, for The abbreviation of , which represents the coordinates in the first coordinate system; is the rotation matrix; is the translation matrix; is the coordinate in the second coordinate system.
[0079] The translation matrix can be expressed as:
[0080] (2)
[0081] in, is the coordinate value of the origin of the second coordinate system in the first coordinate system.
[0082] The rotation matrix can be expressed as:
[0083] (3)
[0084] (4)
[0085] (5)
[0086] (6)
[0087] in, is the heading angle; is the pitch angle; is the flip angle.
[0088] In the embodiment of the present application, since the coordinate origin of the second coordinate system coincides with the coordinate origin of the first coordinate system, the translation matrix is expressed as:
[0089] (7)
[0090] In the first coordinate system, the angle between the emitted ray (i.e., the vertical axis of the different second coordinate system) and the vertical axis of the first coordinate system is the heading angle, which is negative clockwise and positive counterclockwise. Because the operation is performed on the bird's-eye view plane, the pitch angle and roll angle are both 0. Therefore, the rotation matrix can be expressed as:
[0091] (8)
[0092] Furthermore, formulas (8) and (7) can be substituted into the above formula (1) to update formula (1) to obtain the mapping relationship between the first coordinate system and different second coordinate systems. For each second coordinate system, the heading angle is the angle between the longitudinal axis of the second coordinate system and the longitudinal axis of the second coordinate system. Furthermore, the first coordinate information of different sample detection frames in the first coordinate system can be substituted into the updated formula (1) to obtain the second coordinate information of different sample detection frames in each second coordinate system.
[0093] Exemplarily, for each sample detection frame, the first coordinate information of the sample detection frame in the first coordinate system includes the position coordinates of each vertex of the sample detection frame in the first coordinate system. In this case, the position coordinates of each vertex of the sample detection frame in the first coordinate system can be substituted into the updated formula (1) to obtain the position coordinates of each vertex of the sample detection frame in the second coordinate system, that is, the second coordinate information of the sample detection frame in the second coordinate system.
[0094] S203 : Determine sample occlusion ratios corresponding to different sample detection frames according to second coordinate information of different sample detection frames in each second coordinate system.
[0095] The sample occlusion ratio corresponding to each sample detection frame is used to quantify the degree to which the sample detection frame is occluded by other sample detection frames.
[0096] Considering that the same target object may have different degrees of occlusion at different angles, and that the vertical axes of different second coordinate systems are oriented in different directions, the occlusion ratio of the sample detection frame in different second coordinate systems represents the degree of occlusion of the sample detection frame at different angles. Therefore, for each sample detection frame, the occlusion ratio of the sample detection frame in different second coordinate systems can be calculated first, and the average or maximum occlusion ratio of the occlusion ratios in different second coordinate systems can be used as the sample occlusion ratio corresponding to the sample detection frame.
[0097] Furthermore, considering that occlusion is essentially the spatial coverage of background objects by foreground objects, the specific determination needs to be based on the coordinate information of each sample detection frame. Therefore, the occlusion ratio of the sample detection frame in a second coordinate system can be calculated using methods such as the 2D projection area overlap method and the 3D volume overlap method.
[0098] S204 , using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames, the initial model is trained to obtain a target detection model.
[0099] Optionally, sample driving scene images and the sample occlusion ratios corresponding to different sample detection frames can be used as training data. The sample occlusion ratios serve as labels. The initial model is trained using this training data to obtain an object detection model.
[0100] It can be understood that the target detection model has the function of marking the occlusion ratio of the sample detection frame obtained by performing target detection on the sample driving scene graph, that is, the output result of the target detection model includes the detection frame obtained by detection and the occlusion ratio of each detection frame.
[0101] In the above model training method, different second coordinate systems are introduced. These different second coordinate systems share the same origin as the first coordinate system, and the angles between their longitudinal axes and the longitudinal axis of the first coordinate system are different. This means that the different second coordinate systems divide the bird's-eye view plane of the first coordinate system. Based on the mapping relationship between the first coordinate system and the different second coordinate systems, the second coordinate information of different sample detection frames in each second coordinate system can be accurately determined based on the first coordinate information of different sample detection frames in the first coordinate system. Furthermore, based on the second coordinate information of different sample detection frames in each second coordinate system, the sample occlusion ratio corresponding to each sample detection frame is determined. This is equivalent to comprehensively considering the occlusion ratio of the same sample detection frame at different angles. Compared with the solution of manually annotating occlusion ratios, this method avoids the subjectivity and arbitrariness of manual annotation and improves the accuracy of the sample occlusion ratios. In addition, the object detection model obtained by training the initial model using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames can output not only the detection frames but also the occlusion ratios of the detection frames.
[0102] In the embodiment of the present application, the second coordinate information of each sample detection frame in each second coordinate system includes the position coordinates of the vertices of the sample detection frame in each second coordinate system; based on the above embodiment, in order to ensure the accuracy of the sample detection frames corresponding to different sample detection frames, in an exemplary embodiment, Figure 3 As shown, a method for determining the sample occlusion ratio corresponding to different sample detection frames is provided, which refines S203 and specifically includes the following steps:
[0103] S301 : Determine candidate occlusion ratios of different sample detection frames in each second coordinate system according to the position coordinates of the vertices of different sample detection frames in each second coordinate system.
[0104] For each sample detection frame, the candidate occlusion ratio of the sample detection frame in any second coordinate system is the degree to which the sample detection frame is occluded by other sample detection frames in the second coordinate system.
[0105] Optionally, for each sample detection frame, in any of the second coordinate systems, the three-dimensional sample detection frame can be projected into a two-dimensional plane, for example, projected onto a vertical plane of the second coordinate system, and the sum of the intersection areas of the sample detection frame and other sample detection frames is calculated on the vertical plane. The ratio of the sum of the intersection areas to the projection surface of the sample detection frame is used as the candidate occlusion ratio of the sample detection frame in the second coordinate system.
[0106] S302 : For each sample detection frame, taking the maximum value of the candidate occlusion ratios of the sample detection frame in each second coordinate system as the sample occlusion ratio corresponding to the sample detection frame.
[0107] The sample occlusion ratio corresponding to the sample detection frame is the degree to which the sample detection frame is occluded by other sample detection frames in the entire three-dimensional space where the sample vehicle is located.
[0108] Optionally, for each sample detection frame, after calculating the candidate occlusion ratios of the sample detection frame in all second coordinate systems, in order to ensure the accuracy of the sample occlusion ratio corresponding to the determined sample detection frame, the maximum value of the candidate occlusion ratios of the sample detection frame in each second coordinate system can be used as the sample occlusion ratio corresponding to the sample detection frame.
[0109] In this embodiment, for different sample detection frames, the accuracy and availability of the sample occlusion ratio are guaranteed by introducing the candidate occlusion ratios of different sample detection frames in each second coordinate system, and taking the maximum value of the candidate occlusion ratios of the sample detection frame in each second coordinate system as the sample occlusion ratio corresponding to the sample detection frame.
[0110] On the basis of the technical solutions of the above embodiments, in order to achieve accurate calculation of candidate occlusion ratios of different sample detection frames in each second coordinate system, the present application also provides an optional embodiment. In this optional embodiment, Figure 4 As shown, a method for determining a candidate occlusion ratio is provided, which refines S301 and specifically includes the following steps:
[0111] S401 , for each second coordinate system, based on the position coordinates of the vertices of different sample detection frames in the second coordinate system, filter out from each sample detection frame a first detection frame whose vertex vertical coordinates are greater than zero and a second detection frame whose vertex vertical coordinates are less than zero.
[0112] It should be noted that for each second coordinate system, since the coordinate origin of the second coordinate system is the coordinate origin of the vehicle coordinate system, that is, the vehicle center of mass of the sample vehicle, therefore, in this second coordinate system, the sample detection frame in the negative direction of the longitudinal axis will not cause occlusion to the sample detection frame in the positive direction of the longitudinal axis, and the sample detection frame in the negative direction of the longitudinal axis can be regarded as unobstructed in this second coordinate system.
[0113] On this basis, all sample detection frames need to be divided in the second coordinate system, that is, according to the position coordinates of the vertices of different sample detection frames in the second coordinate system, the sample detection frames with vertical coordinates greater than zero in the position coordinates of each vertex can be used as the first detection frame, and the sample detection frames with vertical coordinates less than zero in the position coordinates of each vertex can be used as the second detection frame.
[0114] S402 : Using the lower limit of the occlusion ratio as a candidate occlusion ratio of the second detection frame in the second coordinate system, and determining the candidate occlusion ratio of the first detection frame in the second coordinate system based on the position coordinates of the vertices of the first detection frame in the second coordinate system.
[0115] The lower limit of the occlusion ratio is a preset minimum value of the occlusion ratio, which is generally 0.
[0116] Optionally, since the second detection frame can be regarded as unobstructed in the second coordinate system, the candidate occlusion ratio of the second detection frame in the second coordinate system can be set to the lower limit of the occlusion ratio, for example, can be set to 0.
[0117] In addition, for the first detection frame, if the number of the first detection frame is one, the first detection frame will not be occluded by other first detection frames in the second coordinate system. Therefore, in this case, the candidate occlusion ratio of the first detection frame in the second coordinate system is 0.
[0118] For the first detection frame, if the number of first detection frames is at least two, then for any first detection frame, the three-dimensional sample detection frame can be projected into a two-dimensional plane based on the position coordinates of the vertices of the first detection frame in the second coordinate system, for example, projected to a vertical plane of the second coordinate system, and the sum of the intersection areas of the sample detection frame and other sample detection frames is calculated in the vertical plane, and the ratio between the sum of the intersection areas and the projection surface of the sample detection frame is used as the candidate occlusion ratio of the sample detection frame in the second coordinate system.
[0119] In this embodiment, by dividing different sample detection frames in the same second coordinate system into a first detection frame and a second detection frame, and taking the lower limit value of the occlusion ratio as the candidate occlusion ratio of the second detection frame in the second coordinate system, the calculation amount of the candidate occlusion ratio of the second detection frame is reduced; in addition, the accuracy of determining the candidate occlusion ratio of the first detection frame in the second coordinate system is guaranteed.
[0120] On the basis of the technical solutions of the above embodiments, in order to achieve accurate calculation of the candidate occlusion ratio of the first detection frame in the second coordinate system, the present application also provides an optional embodiment. In this optional embodiment, Figure 5 As shown, a method for determining a candidate occlusion ratio is provided, which refines S402 and specifically includes the following steps:
[0121] S501 : Determine the vertical coordinate of the center point of each first detection frame according to the position coordinates of the vertices of each first detection frame in the second coordinate system.
[0122] Optionally, for any first detection frame, considering that the first detection frame is a three-dimensional detection frame, the number of vertices of the first detection frame is 8. In this case, the vertical coordinate of the center point of the first detection frame is the average value of the vertical coordinates of the 8 vertices.
[0123] S502 , sorting the first detection frames in descending order according to the vertical coordinates of the center points of the first detection frames to obtain serial numbers of the first detection frames.
[0124] It's understandable that the vertical coordinate of the center point of each first detection frame, to some extent, represents the longitudinal distance between the first detection frame and the sample vehicle. Considering that the first detection frame closest to the sample vehicle is unobstructed, and the first detection frame farthest from the sample vehicle is likely to be obstructed, the first detection frames can be sorted in descending order to obtain their sequence numbers. In other words, the first detection frame farthest from the sample vehicle has the smallest sequence number.
[0125] S503 : Determine candidate occlusion ratios of the first detection frames in the second coordinate system in sequence according to the serial numbers of the first detection frames and the position coordinates of the vertices in the second coordinate system.
[0126] For any first detection frame, considering that other first detection frames located after the first detection frame (that is, other first detection frames with serial numbers greater than the serial number of the first detection frame) will not cause occlusion to the first detection frame, it is possible to start from the first detection frame with the smallest serial number and calculate the candidate occlusion ratios of the first detection frames in the second coordinate system in sequence.
[0127] In this embodiment, by sorting each first detection frame according to the vertical coordinate of the center point of each first detection frame, it is equivalent to sorting each first detection frame according to the longitudinal distance between the first detection frame and the sample vehicle. In this case, the candidate occlusion ratio of each first detection frame in the second coordinate system is determined once according to the serial number of each first detection frame, which reduces the amount of calculation and improves the calculation efficiency.
[0128] On the basis of the technical solutions of the above embodiments, in order to achieve accurate calculation of the candidate occlusion ratio of the first detection frame in the second coordinate system, the present application also provides an optional embodiment. In this optional embodiment, Figure 6A As shown, a method for determining a candidate occlusion ratio is provided, which refines S503 and specifically includes the following steps:
[0129] S601 : Projecting each first detection frame onto a vertical plane of the second coordinate system according to the position coordinates of the vertices of each first detection frame in the second coordinate system to obtain a projection surface of each first detection frame on the vertical plane.
[0130] The vertical plane of the second coordinate system is the plane formed by the horizontal and vertical axes of the second coordinate system. The projection plane of each first detection frame on the vertical plane is the projection plane obtained by projecting the first detection frame onto the vertical plane. In this embodiment of the present application, the projection plane of the first detection frame can be represented by the position coordinates (y, z) of each vertex, where y represents the horizontal coordinate of the vertex and z represents the vertical coordinate of the vertex.
[0131] Optionally, for each first detection frame, the vertical coordinates of each vertex of the first detection frame may be forcibly set to 0 to obtain a projection surface of the first detection frame on a vertical plane.
[0132] S602 : For each first detection frame, select a detection frame having a sequence number greater than the sequence number of the first detection frame from among the first detection frames, and use it as a reference detection frame for the first detection frame.
[0133] The reference detection frame of the first detection frame is a detection frame that may block the first detection frame.
[0134] To reduce workload, for each first detection frame, we can first exclude detection frames that are unlikely to obstruct it from all first detection frames. It is understandable that first detection frames with a sequence number smaller than that of the first detection frame are located further away from it and cannot obstruct it. However, other first detection frames with a sequence number larger than that of the first detection frame are located between it and the sample vehicle and may obstruct it.
[0135] Therefore, for each first detection frame, a detection frame with a serial number greater than the serial number of the first detection frame may be screened out from the first detection frames to serve as a reference detection frame for the first detection frame.
[0136] S603: Determine an intersection area between the projection surface of the first detection frame and the projection surface of the reference detection frame.
[0137] The intersection area of the projection surface of the first detection frame and the projection surface of the reference detection frame is the area of the intersection of the projection surface of the first detection frame and the projection surface of the reference detection frame. For example, for two intersecting detection frames K1 and K2, Figure 6B As shown, S1 is the projection area of K1, S2 is the projection area of K2, and the shaded area That is the intersection of the two detection frames, that is, the shadow part The area is the intersection area of projection surface S1 and projection surface S2.
[0138] Optionally, the intersection area between the projection surface of the first detection frame and the projection surface of the reference detection frame can be calculated based on the position coordinates of each vertex of the projection surface of the reference detection frame that intersects the first detection frame, as well as the position coordinates of each vertex of the projection surface of the first detection frame. For reference detection frames that do not intersect with the first detection frame, the corresponding intersection area is 0.
[0139] S604 : Determine a candidate occlusion ratio of the first detection frame in the second coordinate system according to an intersection area between the projection surface of the first detection frame and the projection surface of the reference detection frame, and a projection area of the projection surface of the first detection frame.
[0140] Optionally, the ratio of the sum of the intersection areas of the projection surface of the first detection frame and the projection surface of the reference detection frame to the projection area of the projection surface of the first detection frame can be used as the candidate occlusion ratio of the first detection frame in the second coordinate system.
[0141] For example, if the first detection frame includes K1, K2, K3, and K4, and for K1, the first detection frame with a sequence number greater than K1 includes K2, K3, and K4, then the candidate occlusion ratio of K1 in the second coordinate system can be calculated by the following formula (9):
[0142] (9)
[0143] in, is the candidate occlusion ratio of K1 in the second coordinate system; is the intersection area of K1 and K2; is the intersection area of K1 and K3; is the intersection area of K1 and K4; is the projected area of K1.
[0144] In this embodiment, by projecting the first detection frame onto the vertical plane of the second coordinate system, it is equivalent to constructing a plane in the vertical direction perpendicular to the ray (i.e., the longitudinal axis), which reduces the difficulty of directly constructing the vertical plane of the ray and ensures the accuracy of the candidate occlusion ratio of the determined first detection frame in the second coordinate system.
[0145] In an exemplary embodiment, Figure 7 As shown, a target detection method is provided, which is applied to Figure 1 Taking the vehicle-mounted terminal 102 in FIG. 1 as an example, the method specifically includes the following steps:
[0146] S701, obtaining a current driving scene map of the target vehicle.
[0147] The current driving scene image is an image of the driving environment of the target vehicle collected by the target vehicle at the current moment.
[0148] Optionally, a collection device installed in the target vehicle may collect a driving scene map of the target vehicle's driving environment in real time, and transmit the map to the target vehicle's onboard terminal as the current driving scene map.
[0149] S702 , input the current driving scene image into the target detection model to obtain each initial detection frame output by the target detection model for target detection on the current driving scene image and the target occlusion ratio corresponding to each initial detection frame.
[0150] Optionally, the current driving scene graph can be input into the target detection model. The target detection model will perform target detection on the current driving scene graph based on the trained model parameters, predict the target occlusion ratio of each initial detection frame obtained by the target detection, and output each initial detection frame and the target occlusion ratio corresponding to each initial detection frame.
[0151] S703 : The initial detection frame whose target occlusion ratio among the initial detection frames is less than a preset occlusion threshold is used as the target detection frame.
[0152] The predicted occlusion threshold is a preset threshold of the occlusion ratio, for example, the preset occlusion threshold can be set to 70%. The target detection frame is a valid detection frame obtained by target detection.
[0153] Alternatively, consider that at a certain moment, while target vehicle A is driving normally, other vehicle B is obstructed by a non-motorized vehicle, a pillar, or the like. Therefore, at that moment, other vehicle B will not cross the obstruction and affect the normal driving of target vehicle A. However, if other vehicle B is still used as a detection target at this moment, target vehicle A may mistakenly detect other vehicle B as a target that is about to intrude on target vehicle A's driving path, causing target vehicle A to initiate emergency braking, further affecting the normal driving of target vehicle A. Therefore, initial detection frames with a target occlusion ratio greater than a preset occlusion threshold can be excluded from detection.
[0154] That is to say, only the initial detection frame whose target occlusion ratio is less than the preset occlusion threshold is retained as the target detection frame.
[0155] S704: Determine the target detection result of the target vehicle according to the target detection frame.
[0156] Among them, the target detection results are presented in the form of target detection boxes and category labels, which are used to identify the position and category of the target in the current driving scene image.
[0157] Optionally, the target detection frame includes target objects that will affect the driving process of the target vehicle. Therefore, the position of each target object relative to the target vehicle can be determined based on the detection frame coordinates of the target detection frame, and the detected target objects can be displayed on the display device in the target vehicle. At the same time, the specific category of the target object can be marked so that the driver of the target vehicle can control the target vehicle based on the target detection results.
[0158] The target detection method described above inputs the acquired current driving scene image of the target vehicle into a target detection model, obtains the initial detection frames output by the target detection model for target detection of the current driving scene image, and the target occlusion ratio corresponding to each initial detection frame. The initial detection frames whose target occlusion ratio is less than a preset occlusion threshold are then used as target detection frames, and the target detection results of the target vehicle are determined based on the target detection frames. By introducing a target detection model, the above scheme predicts the target occlusion ratio of each initial detection frame (i.e., each target object) obtained during target detection, ensuring the efficiency and accuracy of determining the target occlusion ratio. The initial detection frames are then screened based on the target occlusion ratio, removing initial detection frames whose target occlusion ratio is greater than the preset occlusion threshold. This prevents target objects that would not otherwise affect the target vehicle's normal driving from affecting its trajectory.
[0159] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0160] Based on the same inventive concept, the present application also provides a model training device for implementing the aforementioned model training method. The solution to the problem provided by the device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following model training device embodiments can be found in the above-mentioned limitations on the model training method, and will not be repeated here.
[0161] In an exemplary embodiment, Figure 8 As shown, a model training device 800 is provided, comprising: a first detection module 810, a coordinate mapping module 820, an occlusion determination module 830 and a model training module 840, wherein:
[0162] The first detection module 810 is used to input the sample driving scene graph of the sample vehicle into the initial model, and obtain the first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene graph in the first coordinate system; the first coordinate system is a vehicle coordinate system with the forward direction of the sample vehicle as the vertical axis.
[0163] The coordinate mapping module 820 is used to determine the second coordinate information of different sample detection frames in each second coordinate system based on the mapping relationship between the first coordinate system and different second coordinate systems and according to the first coordinate information of different sample detection frames in the first coordinate system; each second coordinate system and the first coordinate system have the same coordinate origin, and the angles between the longitudinal axes of different second coordinate systems and the longitudinal axis of the first coordinate system are different.
[0164] The occlusion determination module 830 is configured to determine the sample occlusion ratios corresponding to the different sample detection frames according to the second coordinate information of the different sample detection frames in the respective second coordinate systems.
[0165] The model training module 840 is used to train the initial model using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames to obtain a target detection model.
[0166] The model training device introduces different second coordinate systems, each of which has the same origin as the first coordinate system and has different angles between its longitudinal axis and the longitudinal axis of the first coordinate system. This means that the different second coordinate systems divide the bird's-eye view plane of the first coordinate system. Based on the mapping relationship between the first coordinate system and the different second coordinate systems, the second coordinate information of different sample detection frames in each second coordinate system can be accurately determined based on the first coordinate information of different sample detection frames in the first coordinate system. Furthermore, the sample occlusion ratios corresponding to different sample detection frames are determined based on the second coordinate information of different sample detection frames in each second coordinate system. This is equivalent to comprehensively considering the occlusion ratios of the same sample detection frame at different angles. Compared to manually annotating occlusion ratios, this method avoids the subjectivity and arbitrariness of manual annotation and improves the accuracy of the sample occlusion ratios. Furthermore, the object detection model obtained by training the initial model using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames can output not only the detection frames but also the occlusion ratios of the detection frames.
[0167] In one embodiment, the second coordinate information of each sample detection frame in each second coordinate system includes the position coordinates of the vertices of the sample detection frame in each second coordinate system; the occlusion determination module 830 includes:
[0168] The first determining unit is configured to determine candidate occlusion ratios of different sample detection frames in each second coordinate system according to position coordinates of vertices of different sample detection frames in each second coordinate system.
[0169] The second determining unit is configured to take, for each sample detection frame, a maximum value among the candidate occlusion ratios of the sample detection frame in each second coordinate system as the sample occlusion ratio corresponding to the sample detection frame.
[0170] In one embodiment, the first determining unit includes:
[0171] The first determination slave unit is used to, for each second coordinate system, filter out, from each sample detection frame, a first detection frame whose vertex vertical coordinates are greater than zero and a second detection frame whose vertex vertical coordinates are less than zero, based on the position coordinates of the vertices of different sample detection frames in the second coordinate system.
[0172] The second determining slave unit is used to use the lower limit value of the occlusion ratio as the candidate occlusion ratio of the second detection frame in the second coordinate system; and determine the candidate occlusion ratio of the first detection frame in the second coordinate system according to the position coordinates of the vertices of the first detection frame in the second coordinate system.
[0173] In one embodiment, the second determining slave unit includes:
[0174] The first determining subunit is configured to determine the vertical coordinate of the center point of each first detection frame according to the position coordinates of the vertices of each first detection frame in the second coordinate system.
[0175] The second determining subunit is configured to sort the first detection frames in descending order according to the vertical coordinates of the center points of the first detection frames to obtain serial numbers of the first detection frames.
[0176] The third determining subunit is configured to sequentially determine the candidate occlusion ratio of each first detection frame in the second coordinate system according to the serial number of each first detection frame and the position coordinates of the vertex in the second coordinate system.
[0177] In one embodiment, the third determining subunit is specifically configured to:
[0178] According to the position coordinates of the vertices of each first detection frame in the second coordinate system, each first detection frame is projected onto the vertical plane of the second coordinate system to obtain the projection surface of each first detection frame in the vertical plane; wherein, the vertical plane of the second coordinate system is a plane formed by the horizontal axis and the vertical axis of the second coordinate system; for each first detection frame, a detection frame with a serial number greater than the serial number of the first detection frame is selected from each first detection frame as a reference detection frame of the first detection frame; the intersection area of the projection surface of the first detection frame and the projection surface of the reference detection frame is determined; according to the intersection area of the projection surface of the first detection frame and the projection surface of the reference detection frame, as well as the projection area of the projection surface of the first detection frame, the candidate occlusion ratio of the first detection frame in the second coordinate system is determined.
[0179] In one embodiment, the model training device 800 further includes a coordinate system construction module, specifically configured to:
[0180] In the bird's-eye view plane where the first coordinate system is located, with the coordinate origin of the first coordinate system as the starting point, a preset number of rays are emitted in sequence in the counterclockwise direction; the preset number of rays are used as the vertical axis of each second coordinate system, and the horizontal axis of different second coordinate systems is determined according to the vertical axis of each second coordinate system to construct each second coordinate system.
[0181] In an exemplary embodiment, Figure 9 As shown, a target detection device 900 is provided, comprising: a data acquisition module 910, an occlusion prediction module 920, a data screening module 930 and a second detection module 940, wherein:
[0182] The data acquisition module 910 is used to obtain the current driving scene diagram of the target vehicle.
[0183] The occlusion prediction module 920 is used to input the current driving scene image into the target detection model to obtain the initial detection frames output by the target detection model for target detection on the current driving scene image and the target occlusion ratio corresponding to each initial detection frame.
[0184] The data screening module 930 is configured to select, among the initial detection frames, the initial detection frames whose target occlusion ratio is less than a preset occlusion threshold as target detection frames.
[0185] The second detection module 940 is used to determine the target detection result of the target vehicle according to the target detection frame.
[0186] The target detection device inputs the acquired current driving scene image of the target vehicle into a target detection model, obtains the initial detection frames output by the target detection model for target detection of the current driving scene image, and the target occlusion ratio corresponding to each initial detection frame. The initial detection frames whose target occlusion ratio is less than a preset occlusion threshold are then used as target detection frames, and the target detection results of the target vehicle are determined based on the target detection frames. By introducing a target detection model to predict the target occlusion ratio of each initial detection frame (i.e., each target object) obtained during target detection, the above scheme ensures the efficiency and accuracy of determining the target occlusion ratio. The initial detection frames are then screened based on the target occlusion ratio, and initial detection frames whose target occlusion ratio is greater than the preset occlusion threshold are deleted, thereby preventing target objects that would not otherwise affect the normal driving of the target vehicle from affecting its driving trajectory.
[0187] Each module in the above-mentioned model training device and target detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0188] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 10 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a model training method is implemented.
[0189] In one embodiment, a computer device is provided. The computer device may be a vehicle-mounted terminal, and its internal structure diagram may be as follows: Figure 11 As shown. The computer device includes a processor, memory, communication interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI (Wireless Fidelity), mobile cellular network, NFC (Near Field Communication) or other technologies. When the computer program is executed by the processor, a target detection method is implemented.
[0190] Those skilled in the art will understand that Figure 10 and Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0191] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the model training method provided in the above embodiment is implemented.
[0192] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the target detection method provided in the above embodiment when executing the computer program.
[0193] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the model training method provided in the above embodiment is implemented.
[0194] In one embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the target detection method provided in the above embodiment is implemented.
[0195] In one embodiment, a computer program product is provided, including a computer program, which implements the model training method provided in the above embodiment when executed by a processor.
[0196] In one embodiment, a computer program product is further provided, including a computer program, which implements the target detection method provided in the above embodiment when executed by a processor.
[0197] It should be noted that the data involved in this application (including but not limited to sample driving scene diagrams, current driving scene diagrams, etc.) are information and data that have been fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0198] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0199] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0200] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A model training method, characterized in that: The method comprises: Inputting a sample driving scene graph of a sample vehicle into an initial model, obtaining first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene graph in a first coordinate system; the first coordinate system is a vehicle coordinate system with the sample vehicle's forward direction as the vertical axis; Based on a mapping relationship between the first coordinate system and different second coordinate systems, second coordinate information of different sample detection frames in each second coordinate system is determined according to the first coordinate information of different sample detection frames in the first coordinate system; each second coordinate system and the first coordinate system have the same coordinate origin, and the vertical axes of different second coordinate systems have different angles with the vertical axis of the first coordinate system; Determining the sample occlusion ratios corresponding to the different sample detection frames according to the second coordinate information of the different sample detection frames in the respective second coordinate systems; The sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames are used to train the initial model to obtain a target detection model.
2. The method according to claim 1, characterized in that The second coordinate information of each sample detection frame in each second coordinate system includes the position coordinates of the vertices of the sample detection frame in each second coordinate system; Determining the sample occlusion ratios corresponding to the different sample detection frames according to the second coordinate information of the different sample detection frames in the respective second coordinate systems includes: Determining candidate occlusion ratios of different sample detection frames in each second coordinate system according to the position coordinates of the vertices of different sample detection frames in each second coordinate system; For each sample detection frame, the maximum value among the candidate occlusion ratios of the sample detection frame in each second coordinate system is used as the sample occlusion ratio corresponding to the sample detection frame.
3. The method according to claim 2, characterized in that Determining candidate occlusion ratios of different sample detection frames in each second coordinate system according to the position coordinates of the vertices of different sample detection frames in each second coordinate system includes: For each second coordinate system, based on the position coordinates of the vertices of different sample detection frames in the second coordinate system, filter out from each sample detection frame a first detection frame whose vertex vertical coordinates are greater than zero and a second detection frame whose vertex vertical coordinates are less than zero; Using the lower limit of the occlusion ratio as a candidate occlusion ratio of the second detection frame in the second coordinate system; and Determine candidate occlusion ratios of the first detection frame in the second coordinate system according to position coordinates of vertices of the first detection frame in the second coordinate system.
4. The method according to claim 3, characterized in that When the number of the first detection frames is at least two, determining, according to position coordinates of vertices of the first detection frames in the second coordinate system, candidate occlusion ratios of the first detection frames in the second coordinate system includes: determining the vertical coordinate of the center point of each first detection frame according to the position coordinates of the vertices of each first detection frame in the second coordinate system; Sort the first detection frames in descending order according to the vertical coordinates of the center points of the first detection frames to obtain serial numbers of the first detection frames; According to the serial number of each first detection frame and the position coordinates of the vertex in the second coordinate system, the candidate occlusion ratio of each first detection frame in the second coordinate system is determined in sequence.
5. The method according to claim 4, characterized in that Determining, in sequence, the candidate occlusion ratios of the first detection frames in the second coordinate system according to the serial numbers of the first detection frames and the position coordinates of the vertices in the second coordinate system, including: Projecting each first detection frame onto a vertical plane of the second coordinate system based on the position coordinates of the vertices of each first detection frame in the second coordinate system to obtain a projection plane of each first detection frame on the vertical plane; wherein the vertical plane of the second coordinate system is a plane formed by the horizontal axis and the vertical axis of the second coordinate system; For each first detection frame, select a detection frame with a serial number greater than the serial number of the first detection frame from the first detection frames to serve as a reference detection frame for the first detection frame; Determining an intersection area between a projection surface of the first detection frame and a projection surface of the reference detection frame; Determine a candidate occlusion ratio of the first detection frame in the second coordinate system according to an intersection area of a projection surface of the first detection frame and a projection surface of the reference detection frame, and a projection area of the projection surface of the first detection frame.
6. The method according to any one of claims 1 to 5, characterized in that Different second coordinate systems are constructed in the following ways: In the bird's-eye view plane where the first coordinate system is located, a preset number of rays are emitted in sequence in a counterclockwise direction with the coordinate origin of the first coordinate system as the starting point; A preset number of rays are used as the vertical axis of each second coordinate system, and the horizontal axes of different second coordinate systems are determined according to the vertical axis of each second coordinate system to construct each second coordinate system.
7. A target detection method, characterized in that: The method comprises: Obtain the current driving scene graph of the target vehicle; Inputting the current driving scene image into a target detection model to obtain initial detection frames output by the target detection model for target detection of the current driving scene image and target occlusion ratios corresponding to each initial detection frame; wherein the target detection model is trained based on the model training method according to any one of claims 1 to 6; The initial detection frame whose target occlusion ratio in each initial detection frame is less than the preset occlusion threshold is used as the target detection frame; A target detection result of the target vehicle is determined according to the target detection frame.
8. A model training device, characterized in that: The device comprises: a first detection module, configured to input a sample driving scene graph of a sample vehicle into an initial model, and obtain first coordinate information of different sample detection frames output by the initial model for target detection on the sample driving scene graph in a first coordinate system; the first coordinate system being a vehicle coordinate system with the sample vehicle's forward direction as its vertical axis; a coordinate mapping module, configured to determine, based on a mapping relationship between the first coordinate system and different second coordinate systems, second coordinate information of different sample detection frames in each second coordinate system according to the first coordinate information of different sample detection frames in the first coordinate system; each second coordinate system and the first coordinate system have the same coordinate origin, and the vertical axes of different second coordinate systems have different angles with the vertical axis of the first coordinate system; an occlusion determination module, configured to determine the sample occlusion ratios corresponding to different sample detection frames based on the second coordinate information of the different sample detection frames in each second coordinate system; The model training module is used to train the initial model using the sample driving scene graph and the sample occlusion ratios corresponding to different sample detection frames to obtain a target detection model.
9. A target detection device, characterized in that: The device comprises: A data acquisition module is used to obtain the current driving scene map of the target vehicle; an occlusion prediction module, configured to input the current driving scene image into a target detection model, and obtain initial detection frames output by the target detection model for target detection of the current driving scene image and a target occlusion ratio corresponding to each initial detection frame; wherein the target detection model is trained based on the model training method according to any one of claims 1 to 6; A data screening module is used to select the initial detection frames whose target occlusion ratio in each initial detection frame is less than a preset occlusion threshold as the target detection frame; The second detection module is used to determine the target detection result of the target vehicle according to the target detection frame.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.