Method for obtaining wheel contact point, storage medium and electronic device
By identifying the center point of the wheel touchdown and image substructure in the candidate recognition image, the insufficient acquisition of wheel touchdown information is solved, and efficient and accurate wheel touchdown recognition is achieved on low-computing equipment, supporting automatic driving and accurate position detection of ADAS.
Patent Information
- Application Number
- CN202210232479.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-03-09
AI Technical Summary
The lack of effective way to obtain wheel touch point information in the prior art, resulting in the inability to directly apply to detection tasks.
By identifying the center point of the wheel touchdown of the target vehicle in the candidate recognition image, the image substructure of the candidate recognition image is obtained, and the position information of the wheel touchdown point is determined by using these substructures, and the recognition efficiency is improved in combination with the anchor point mechanism.
It realizes efficient acquisition of wheel touch point information on low-computing equipment, improves the accuracy and comprehensiveness of wheel touch point identification, and supports accurate position detection of autonomous driving and ADAS.
Smart Images

Figure CN114627069B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a method for obtaining a wheel contact point, a storage medium, and an electronic device. Background Art
[0002] Wheel contact point detection has become increasingly common in recent years, but current methods for acquiring this information are still immature and face various challenges. Consequently, there's no way to directly use this information for detection. Consequently, there's a lack of an effective method for acquiring this information.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a method for obtaining a wheel contact point, a storage medium, and an electronic device to at least solve the technical problem of the lack of an effective method for obtaining wheel contact point information.
[0005] According to one aspect of an embodiment of the present application, a method for obtaining a wheel contact point is provided, comprising: upon obtaining a candidate recognition image and identifying at least one target vehicle from the candidate recognition image, determining the wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point is used to represent position information of centers of at least two wheel contact points of the target vehicle in the candidate recognition image; obtaining at least two image substructures corresponding to the candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to obtain position information of the wheel contact point; determining a target image substructure corresponding to each target vehicle from the at least two image substructures based on the image area where the wheel contact center point is located; and obtaining position information of the wheel contact point of each target vehicle using the target image substructures.
[0006] According to another aspect of an embodiment of the present application, a device for acquiring a wheel contact point is further provided, comprising a first determination unit for determining, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, a wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point is used to represent position information of centers of at least two wheel contact points of the target vehicle in the candidate recognition image; a first acquisition unit for acquiring at least two image substructures corresponding to the candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to acquire position information of the wheel contact point; a second determination unit for determining, from the at least two image substructures, a target image substructure corresponding to each target vehicle based on the image area where the wheel contact center point is located; and a second acquisition unit for acquiring position information of the wheel contact point of each target vehicle using the target image substructure.
[0007] As an optional solution, the above-mentioned second determination unit includes: a first acquisition module, used to obtain a set of image areas corresponding to the above-mentioned at least two image substructures, wherein the above-mentioned image area set includes at least two of the above-mentioned image areas in the above-mentioned candidate recognition image; a second acquisition module, used to obtain the target image area where the wheel contact center point corresponding to each of the above-mentioned target vehicles is located, wherein the above-mentioned at least two image areas include the above-mentioned target image area; and a first determination module, used to determine the above-mentioned target image substructure corresponding to each of the above-mentioned target image areas from the above-mentioned at least two image substructures.
[0008] As an optional solution, the above-mentioned first acquisition module includes: a first acquisition submodule, used to acquire a first area set corresponding to at least two first image substructures, wherein the above-mentioned first area set includes at least two first image areas in the above-mentioned candidate recognition image, and the above-mentioned first image areas correspond to the above-mentioned first image substructure; a second acquisition submodule, used to acquire a second area set corresponding to at least two second image substructures, wherein the above-mentioned second area set includes at least two second image areas in the above-mentioned candidate recognition image, the above-mentioned second image areas correspond to the above-mentioned second image substructure, and the area range of the above-mentioned second image area is larger than the area range of the above-mentioned first image area.
[0009] As an optional solution, the above-mentioned first determination module includes: a first determination submodule, used to determine the target first substructure and the target second substructure corresponding to each of the above-mentioned target image areas, wherein the above-mentioned target image substructure includes the above-mentioned target first substructure and the above-mentioned target second substructure; or, a second determination submodule, used to determine the above-mentioned target first substructure corresponding to each of the above-mentioned target image areas when the space occupancy of the above-mentioned target vehicle is less than the first threshold, and to determine the above-mentioned target second substructure corresponding to each of the above-mentioned target image areas when the space occupancy of the above-mentioned target vehicle is greater than or equal to the above-mentioned first threshold.
[0010] As an optional solution, the above-mentioned second acquisition unit includes: a third acquisition module, used to obtain the regional center point of each of the above-mentioned target image areas; a calculation module, used to use the above-mentioned target image substructure to calculate the target offset between each of the above-mentioned regional center points and the above-mentioned wheel contact point of each corresponding above-mentioned target vehicle; and a fourth acquisition module, used to obtain the position information of the above-mentioned wheel contact point of each above-mentioned target vehicle based on each of the above-mentioned target offsets and the regional attribute information of the above-mentioned target image area.
[0011] As an optional solution, the above-mentioned device also includes: an input unit, used to input the above-mentioned candidate recognition image into an image recognition model, wherein the above-mentioned image recognition model is a neural network model obtained after training using multiple sample image data and used to identify the position information of the above-mentioned wheel contact points in the image; a third acquisition unit, used to obtain the recognition result output by the above-mentioned image recognition model, wherein the above-mentioned recognition result includes the position information of the above-mentioned wheel contact points of each of the above-mentioned target vehicles in the above-mentioned candidate recognition image.
[0012] As an optional solution, it includes: a fourth acquisition unit, used to acquire the above-mentioned multiple sample image data before inputting the above-mentioned candidate recognition image into the image recognition model; a first marking unit, used to mark the image data used to represent the above-mentioned wheel contact point in each of the above-mentioned sample image data before inputting the above-mentioned candidate recognition image into the image recognition model, so as to obtain the above-mentioned multiple marked sample image data, wherein each marked sample image data includes a marked wheel contact point identifier; a first training unit, used to input the above-mentioned multiple marked sample image data into the initial image recognition model before inputting the above-mentioned candidate recognition image into the image recognition model, so as to train the above-mentioned image recognition model.
[0013] As an optional solution, the above-mentioned first training unit includes: a first repetition module, which is used to repeatedly execute the following steps until the above-mentioned image recognition model is obtained; a second determination module, which is used to determine the current sample image data from the above-mentioned multiple sample image data after labeling, and determine the current image recognition model, wherein the above-mentioned current sample image data includes the labeled current wheel contact point identifier; a first recognition module, which is used to identify the first result data representing the center point of the wheel contact in the above-mentioned current sample image data through the above-mentioned current image recognition model; a first processing module, which is used to process the above-mentioned first result data through the above-mentioned current image recognition model to obtain the second result data representing the position information of the above-mentioned wheel contact point; a second processing module, which is used to obtain the next sample image data as the above-mentioned current sample image data when the above-mentioned second result data does not meet the recognition convergence condition; a third processing module, which is used to determine that the above-mentioned current image recognition model is the above-mentioned image recognition model when the above-mentioned second result data meets the above-mentioned recognition convergence condition.
[0014] As an optional solution, it includes: a fifth acquisition unit, used to acquire the above-mentioned multiple sample image data before inputting the above-mentioned candidate recognition image into the image recognition model; a second marking unit, used to mark the first image data used to represent the above-mentioned wheel contact point and the second image data used to represent the above-mentioned wheel contact center point in each of the above-mentioned sample image data before inputting the above-mentioned candidate recognition image into the image recognition model, so as to obtain the above-mentioned multiple marked sample image data, wherein each marked sample image data includes a marked wheel contact point identifier and a marked wheel contact center point identifier; a second training unit, used to input the above-mentioned multiple marked sample image data into the initial image recognition model before inputting the above-mentioned candidate recognition image into the image recognition model, so as to train the above-mentioned image recognition model.
[0015] As an optional solution, the above-mentioned second training unit includes: a second repetition module, which is used to repeatedly execute the following steps until the above-mentioned image recognition model is obtained; a third determination module, which is used to determine the current sample image data from the above-mentioned multiple labeled sample image data and determine the current image recognition model, wherein the above-mentioned current sample image data includes the labeled current wheel contact point identifier; a second recognition module, which is used to identify the first result data representing the center point of the wheel contact in the above-mentioned current sample image data through the above-mentioned current image recognition model; a fourth processing module, which is used to obtain the next sample image data as the above-mentioned current sample image data when the above-mentioned first result data does not meet the second convergence condition; a fifth processing module, which is used to process the above-mentioned first result data through the above-mentioned current image recognition model when the above-mentioned first result data meets the above-mentioned second convergence condition to obtain the second result data representing the position information of the above-mentioned wheel contact point; a sixth processing module, which is used to obtain the next sample image data as the above-mentioned current sample image data when the above-mentioned second result data does not meet the second convergence condition; and a seventh processing module, which is used to determine that the above-mentioned current image recognition model is the above-mentioned image recognition model when the above-mentioned second result data meets the above-mentioned second convergence condition.
[0016] As an optional solution, the above-mentioned second acquisition unit includes at least one of the following: a first position module, used to use the above-mentioned target image substructure to obtain the first position information of the above-mentioned wheel contact point visible to each of the above-mentioned target vehicles; a second position module, used to use the above-mentioned target image substructure to obtain the second position information of the above-mentioned wheel contact point invisible to each of the above-mentioned target vehicles.
[0017] As an optional solution, the above-mentioned second acquisition unit includes: a fifth acquisition module, which is used to use the above-mentioned target image substructure to obtain the contact point set corresponding to each of the above-mentioned target vehicles, wherein the above-mentioned contact point set includes multiple wheel contact points; the above-mentioned device includes: after using the above-mentioned target image substructure to obtain the position information of the above-mentioned wheel contact points of each of the above-mentioned target vehicles, determining a target contact point set from the contact point set corresponding to each of the above-mentioned target vehicles, wherein the number of the above-mentioned wheel contact points in the above-mentioned target contact point set is greater than a second threshold; and using a non-maximum suppression method to filter the above-mentioned wheel contact points in the above-mentioned target contact point set, wherein the number of the above-mentioned wheel contact points in the above-mentioned target contact point set after filtering is less than or equal to the above-mentioned second threshold.
[0018] According to another aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method for obtaining a wheel contact point.
[0019] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method for obtaining the wheel contact point through the computer program.
[0020] In an embodiment of the present application, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, the wheel contact center point of each target vehicle is determined based on the image features of the candidate recognition image, wherein the wheel contact center point is used to represent the position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image; at least two image substructures corresponding to the candidate recognition image are acquired, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, and the image substructure is used to identify the image content information of the candidate recognition image to obtain the position information of the wheel contact point; and according to the image area where the wheel contact center point is located, the wheel contact center point of each target vehicle is determined from the at least two image substructures. From the corresponding target image substructure; using the above-mentioned target image substructure to obtain the position information of the above-mentioned wheel contact point of each above-mentioned target vehicle; the embodiment of the present application uses the target image substructure to obtain the position information of the wheel contact point of each target vehicle, and uses the anchor point mechanism to determine the target image substructure corresponding to each vehicle through the wheel contact center point, and then uses the target image substructure to perform more targeted wheel contact point identification on the corresponding vehicle, and the position information of the identified wheel contact point also establishes an information association relationship with the corresponding vehicle, making the position information of the wheel contact point more comprehensive; in addition, using the wheel contact center point to determine the target image substructure corresponding to each target vehicle from multiple image substructures, improves the application efficiency of the image substructure, and thus solves the technical problem of the lack of an effective method for obtaining wheel contact point information. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0022] Figure 1 is a schematic diagram of an application environment of an optional method for obtaining a wheel contact point according to an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of a process of an optional method for obtaining a wheel contact point according to an embodiment of the present application;
[0024] Figure 3 is a schematic diagram of an optional method for obtaining a wheel contact point according to an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of another optional method for obtaining the wheel contact point according to an embodiment of the present application;
[0026] Figure 5 is a schematic diagram of another optional method for obtaining the wheel contact point according to an embodiment of the present application;
[0027] Figure 6 is a schematic diagram of another optional method for obtaining the wheel contact point according to an embodiment of the present application;
[0028] Figure 7 is a schematic diagram of another optional method for obtaining the wheel contact point according to an embodiment of the present application;
[0029] Figure 8 is a schematic diagram of another optional method for obtaining the wheel contact point according to an embodiment of the present application;
[0030] Figure 9 is a schematic diagram of an optional device for obtaining a wheel contact point according to an embodiment of the present application;
[0031] Figure 10 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] According to one aspect of the embodiment of the present application, a method for obtaining a wheel contact point is provided. Optionally, as an optional implementation, the above-mentioned method for obtaining a wheel contact point can be applied to, but is not limited to, Figure 1 In the environment shown, it may include, but is not limited to, a user device 102, a network 110, and a server 112. The user device 102 may include, but is not limited to, a display 108, a processor 106, and a memory 104. Specifically, the user device 102 may be, but is not limited to, understood as a vehicle-mounted camera for photographing a vehicle, and the vehicle-mounted camera includes a vehicle-mounted camera display (display 108) and a vehicle-mounted camera lens for capturing images.
[0035] The specific process can be as follows:
[0036] Step S102 , the user device 102 acquires a candidate recognition image containing a target vehicle through a vehicle-mounted camera lens;
[0037] Steps S104-S106, the user device 102 sends the candidate recognition image to the server 112 via the network 110;
[0038] In step S108, the server 112 searches the database 114 for at least two image substructures corresponding to the candidate recognition image, and uses the processing engine 116 to determine a target image substructure corresponding to each target vehicle from the at least two image substructures. The server 112 then processes the candidate recognition image using the target image substructure to obtain the location information of the wheel contact point of each target vehicle.
[0039] In steps S110 - S112 , the server 112 sends the location information of the wheel contact point to the user device 102 via the network 110 . The processor 106 in the user device 102 displays the location information of the wheel contact point on the display 108 and stores the location information of the wheel contact point in the memory 104 .
[0040] remove Figure 1 In addition to the examples shown, the above steps can be independently completed by user device 102. Specifically, user device 102 can perform steps such as image processing and obtaining location information of the wheel contact point, thereby reducing the processing pressure on the server. User device 102 includes but is not limited to handheld devices (such as mobile phones), laptops, desktop computers, and in-vehicle devices. This application does not limit the specific implementation of user device 102.
[0041] Alternatively, as an optional implementation, as Figure 2 As shown, the method for obtaining the wheel contact point includes:
[0042] S202, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, determining the wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point is used to represent position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image;
[0043] S204, obtaining at least two image substructures corresponding to the candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image region in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to obtain position information of the wheel contact point;
[0044] S206, determining a target image substructure corresponding to each target vehicle from the at least two image substructures according to the image region where the wheel contact center point is located;
[0045] S208 , using the target image substructure to obtain position information of the wheel contact point of each target vehicle.
[0046] Optionally, in this embodiment, the above-mentioned wheel contact point acquisition method can be applied to, but is not limited to, real-time end-to-end wheel contact point detection scenarios. For example, a detection model (image substructure) is constructed based on an ultra-lightweight convolutional neural network, and a key point (wheel contact center point)-based method is used to detect the wheel contact point, supporting the prediction of occluded contact points. Moreover, due to the use of a structural selection method based on an anchor point mechanism to determine the detection model (target image substructure) corresponding to each target vehicle, the processing efficiency of the detection model is improved. For example, a target vehicle only needs to be processed using part of the detection submodel in the overall detection model. In this way, deployment can be completed on a CPU device with low computing power (such as a vehicle computer), and further real-time detection of the wheel contact points of other vehicles in the current driving environment can be achieved.
[0047] Optionally, in this embodiment, the wheel contact point can, but is not limited to, help the autonomous vehicle obtain a more accurate relative position, direction, and size of the vehicle, which is of great significance to improving safety during driving; for example, after using the target image substructure to obtain the position information of the wheel contact point of each target vehicle, the position information of the wheel contact point is used to determine the spatial information of the target vehicle, wherein the spatial information can, but is not limited to, include at least one of the following: vehicle relative position, vehicle relative direction, vehicle size, vehicle driving trajectory, etc.
[0048] Optionally, in this embodiment, the above-mentioned method for obtaining the wheel contact point can also be used, but is not limited to, in scenarios where more accurate relative distance detection of other vehicles is required. For example, based on the obtained contact point, its corresponding center (pixel position) is calculated, and then converted to the world coordinate position in the vehicle body coordinate system through pre-calibrated camera parameters and ground assumptions. This can help ADAS warnings, lane-level vehicle rendering, etc. obtain accurate position information of other vehicles.
[0049] Optionally, in this embodiment, the above-mentioned method for obtaining the wheel contact point can also be used, but is not limited to, in scenarios where a monocular camera performs 3D target detection of a vehicle. For example, based on the obtained contact point, combined with 2D target detection and camera calibration parameters, 3D target detection information of the vehicle can be generated, including the center position (x, y, z), length, width, height, and heading angle.
[0050] Optionally, in this embodiment, the method of acquiring the candidate recognition image may be, but is not limited to, a method of acquiring images of the vehicle during operation using an on-board image acquisition device, or a method of acquiring road condition images using a roadside camera.
[0051] Optionally, in this embodiment, image features can be understood as, but not limited to, features presented by image pixels in the candidate recognition image, or can be understood as features identified by image recognition technology, such as color features, shape features, local image features, image depth features, edge features, linear features, texture features, etc.; determining the wheel contact center point of each target vehicle based on the image features of the candidate recognition image can be accomplished by, but not limited to, using an image recognition model, for example, first determining an initial image recognition model, and then inputting labeled sample image data for training to obtain a trained image recognition model, wherein the labeled sample image data includes image data of the target vehicle with labeled key points, and the key points include wheel contact points, and specific attribute data of the wheel contact points, such as the left front wheel, left rear wheel, right front wheel and right rear wheel, or visible wheel contact points, invisible wheel contact points, etc.
[0052] Optionally, in this embodiment, the wheel contact center point can be understood as, but not limited to, the center of at least two wheel contact points of the target vehicle. Since the at least two wheel contact points may include a currently invisible wheel contact point, in this case, it is possible but not limited to first predicting the position information of the invisible wheel contact points, and then combining the position information of the visible wheel contact points to predict the position information of the wheel contact center point.
[0053] Optionally, in this embodiment, the target image model may include, but is not limited to, at least two image substructures, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, or it can be understood that each image substructure corresponds to at least one image area in the candidate recognition image, and thus the corresponding image substructure can be directly determined when the image area is determined first, and vice versa; in addition, the image substructure is used to identify the image content information of the candidate recognition image to obtain the position information of the wheel contact point, or it can be understood that the image substructure corresponds to the image area, but the image substructure is used to identify the overall image content information of the candidate recognition image, or is not limited to the image content information of the image area corresponding to the image substructure, but can be, but is not limited to, a stronger correlation with the image content information of the corresponding image area, or the recognition accuracy of the image content information of the corresponding image area is higher, thereby improving the recognition accuracy of the wheel contact point.
[0054] Optionally, in this embodiment, based on the image region where the wheel contact center is located, a target image substructure corresponding to each target vehicle is determined from at least two image substructures. Alternatively, this can be understood as determining the target image region where the wheel contact center is located from multiple image regions, and then using the correspondence between the image region and the image substructures, determining the target image substructure corresponding to the target image region from at least two image substructures. Furthermore, since the target image region is determined by the wheel contact center, and the wheel contact center clearly belongs to a target vehicle, when the target image substructure is determined, the target vehicle corresponding to the target image substructure can also be obtained. Alternatively, this can be understood as a correspondence between the target image substructure and the target vehicle.
[0055] Optionally, in this embodiment, the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle. For example, the target image area in the candidate recognition image is first determined by using the correspondence between the target image substructure and the image area, and then the area center point of the target image area is determined; further, the anchor point mechanism is used to calculate the position offset between the wheel contact point of each target vehicle and the area center point using the target image substructure, and then the actual position information of the wheel contact point of each target vehicle is calculated based on the position offset.
[0056] It should be noted that the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle, and the anchor point mechanism is used to determine the target image substructure corresponding to each vehicle through the wheel contact center point. The target image substructure is then used to perform more targeted wheel contact point identification on the corresponding vehicle, and the position information of the identified wheel contact point is also associated with the corresponding vehicle, making the position information of the wheel contact point more comprehensive. In addition, the wheel contact center point is used to determine the target image substructure corresponding to each target vehicle from multiple image substructures, thereby improving the application efficiency of the image substructure. In this way, the above-mentioned wheel contact point acquisition method can be implemented in application scenarios with lower computing power, such as application scenarios of vehicle-mounted equipment.
[0057] To illustrate further, the optional Figure 3 As shown, when a candidate recognition image 302 is acquired and at least one target vehicle (such as target vehicle 304-1 and target vehicle 304-2) is identified from the candidate recognition image 302, the wheel contact center point (such as wheel contact center point 306-1 and wheel contact center point 306-2) of each target vehicle is determined based on the image features of the candidate recognition image 302, where the wheel contact center point is used to represent the position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image 302;
[0058] Furthermore, at least two image substructures (e.g., image substructure 308-1, image substructure 308-2, and image substructure 308-3) corresponding to the candidate recognition image 302 are obtained, wherein each of the at least two image substructures corresponds to at least one image region in the candidate recognition image 302, and the image substructures are used to identify image content information of the candidate recognition image 302 to obtain position information of the wheel contact point. Based on the image region where the wheel contact center point is located, a target image substructure corresponding to each target vehicle is determined from the at least two image substructures, e.g., the wheel contact center point 306-2 corresponds to the image substructure 308-1, and the wheel contact center point 306-1 corresponds to the image substructure 308-3.
[0059] The target image substructure is further used to obtain the position information of the wheel contact point of each target vehicle, such as using the image substructure 308-1 to obtain the position information of the wheel contact point 310-2 of the target vehicle 304-2, and using the image substructure 308-3 to obtain the position information of the wheel contact point 310-1 of the target vehicle 304-1. In addition, in order to improve the comprehensiveness of the position information of the wheel contact point, additional attribute information can also be provided, such as but not limited to. Figure 3 The solid circle shown in represents the visible wheel contact point, and the dotted circle represents the invisible wheel contact point.
[0060] Through the embodiments provided by the present application, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, the wheel contact center point of each target vehicle is determined based on the image features of the candidate recognition image, wherein the wheel contact center point is used to represent the position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image; at least two image substructures corresponding to the candidate recognition image are acquired, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, and the image substructure is used to identify the image content information of the candidate recognition image to obtain the position information of the wheel contact point; and each target vehicle is determined from the at least two image substructures based on the image area where the wheel contact center point is located. The target image substructure corresponding to each target vehicle is used; the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle; the embodiment of the present application uses the target image substructure to obtain the position information of the wheel contact point of each target vehicle, uses the anchor point mechanism, and determines the target image substructure corresponding to each vehicle through the wheel contact center point, and then uses the target image substructure to perform more targeted wheel contact point identification for each corresponding vehicle, and the position information of the identified wheel contact point is also associated with the corresponding vehicle, making the position information of the wheel contact point more comprehensive; in addition, the wheel contact center point is used to determine the target image substructure corresponding to each target vehicle from multiple image substructures, thereby improving the application efficiency of the image substructure.
[0061] As an optional solution, based on the image area where the wheel contact center point is located, determining the target image substructure corresponding to each target vehicle from at least two image substructures includes:
[0062] S1, obtaining an image region set corresponding to at least two image substructures, wherein the image region set includes at least two image regions in a candidate recognition image;
[0063] S2, obtaining a target image region where the wheel contact center point corresponding to each target vehicle is located, wherein at least two image regions include the target image region;
[0064] S3: Determine a target image substructure corresponding to each target image region from the at least two image substructures.
[0065] Optionally, in this embodiment, based on the image region where the wheel contact center is located, a target image substructure corresponding to each target vehicle is determined from at least two image substructures. Alternatively, this can be understood as determining the target image region where the wheel contact center is located from multiple image regions, and then using the correspondence between the image region and the image substructures, determining the target image substructure corresponding to the target image region from at least two image substructures. Furthermore, since the target image region is determined by the wheel contact center, and the wheel contact center clearly belongs to a target vehicle, when the target image substructure is determined, the target vehicle corresponding to the target image substructure can also be obtained. Alternatively, this can be understood as a correspondence between the target image substructure and the target vehicle.
[0066] To illustrate further, the optional Figure 4 As shown, a set of image regions corresponding to at least two image substructures is obtained, wherein the image region set includes at least two image regions (e.g., a plurality of grid-like structures of 12) in the candidate recognition image 402. A target image region corresponding to the wheel contact center point of each target vehicle is obtained, for example, a target image region 408 where the vehicle contact center point 406 is located is determined from the at least two image regions. Then, using the correspondence between the image regions and the image substructures, a target image substructure corresponding to each target image region 408 is determined from the at least two image substructures.
[0067] Through the embodiments provided in the present application, a set of image regions corresponding to at least two image substructures is obtained, wherein the set of image regions includes at least two image regions in the candidate recognition image; a target image region where the wheel contact center point corresponding to each target vehicle is located is obtained, wherein the at least two image regions include the target image region; and a target image substructure corresponding to each target image region is determined from the at least two image substructures, thereby improving the efficiency of acquiring the target image substructures.
[0068] As an optional solution, obtaining a set of image regions corresponding to at least two image substructures includes:
[0069] S1, obtaining a first region set corresponding to at least two first image substructures, wherein the first region set includes at least two first image regions in the candidate recognition image, and the first image regions correspond to the first image substructures;
[0070] S2. Obtain a second region set corresponding to at least two second image substructures, wherein the second region set includes at least two second image regions in the candidate recognition image, the second image regions correspond to the second image substructures, and the area range of the second image regions is larger than the area range of the first image region.
[0071] Optionally, in this embodiment, the image substructure may be, but is not limited to, an image structure that is divided into multiple categories. This is because there are many types of target vehicles to be identified, and corresponding image structures are selected for targeted processing based on the different types of target vehicles. For example, if the target vehicle is identified as type A, the image substructure corresponding to type A is used to process the target vehicle.
[0072] In addition, in order to improve the efficiency of image processing, the image substructures can be distinguished by, but not limited to, the area range of the image area. For example, the image substructure corresponding to the image area with a smaller area range is used to process the image of the small car, and the image substructure corresponding to the image area with a larger area range is used to process the image of the large car.
[0073] It should be noted that, a first area set corresponding to at least two first image substructures is obtained, wherein the first area set includes at least two first image areas in the candidate recognition image, and the first image areas correspond to the first image substructure; a second area set corresponding to at least two second image substructures is obtained, wherein the second area set includes at least two second image areas in the candidate recognition image, the second image areas correspond to the second image substructure, and the area range of the second image areas is larger than the area range of the first image areas.
[0074] To further illustrate, the optional Figure 4 The scenario shown, continuing with e.g. Figure 5 As shown, a first region set corresponding to at least two first image substructures is obtained, wherein the first region set includes at least two first image regions (a grid with a smaller region range, the number of which is 4×8) in the candidate recognition image, and the first image regions correspond to the first image substructures; further, a target image region 502 where the wheel contact center point 406 is located is determined from the at least two first image regions;
[0075] Optional e.g. Figure 4As shown, a second region set corresponding to at least two second image substructures is obtained, wherein the second region set includes at least two second image regions (grids with smaller region ranges, the number of which is 3×4) in the candidate recognition image, the second image regions corresponding to the second image substructures, and the region range of the second image regions is larger than the region range of the first image regions; further, a target image region 408 where the wheel contact center point 406 is located is determined from the at least two first image regions.
[0076] Through the embodiments provided in the present application, a first area set corresponding to at least two first image substructures is obtained, wherein the first area set includes at least two first image areas in the candidate recognition image, and the first image areas correspond to the first image substructure; a second area set corresponding to at least two second image substructures is obtained, wherein the second area set includes at least two second image areas in the candidate recognition image, and the second image areas correspond to the second image substructure, and the area range of the second image area is larger than the area range of the first image area, thereby achieving the effect of improving the accuracy of the position information of the wheel contact point.
[0077] As an optional solution, determining the target image substructure corresponding to each target image region from the at least two image substructures includes:
[0078] S1, determining a target first substructure and a target second substructure corresponding to each target image region, wherein the target image substructure includes the target first substructure and the target second substructure; or
[0079] S2, when the space occupied by the target vehicle is less than the first threshold, determine the target first substructure corresponding to each target image area; and when the space occupied by the target vehicle is greater than or equal to the first threshold, determine the target second substructure corresponding to each target image area.
[0080] Optionally, in this embodiment, in order to improve the image processing efficiency, the image substructures can be distinguished by, but not limited to, the area range of the image area. For example, the image substructure corresponding to the image area with a smaller area range is used to perform image processing on a small car (the space occupancy is less than a first threshold), and the image substructure corresponding to the image area with a larger area range is used to perform image processing on a large car (the space occupancy is greater than or equal to the first threshold).
[0081] As an optional solution, the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle, including:
[0082] S1, obtain the center point of each target image area;
[0083] S2, using the target image substructure to calculate the target offset between the center point of each region and the wheel contact point of the corresponding target vehicle;
[0084] S3, obtaining the position information of the wheel contact point of each target vehicle according to each target offset and the region attribute information of the corresponding target image region.
[0085] Optionally, in this embodiment, the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle. For example, the target image area in the candidate recognition image is first determined by using the correspondence between the target image substructure and the image area, and then the area center point of the target image area is determined; further, the anchor point mechanism is used to calculate the position offset between the wheel contact point of each target vehicle and the area center point using the target image substructure, and then the actual position information of the wheel contact point of each target vehicle is calculated based on the position offset.
[0086] Optionally, in this embodiment, the target offset may be understood as, but not limited to, the position distance between the center point of each area and the wheel contact point of the corresponding target vehicle.
[0087] Optionally, in this embodiment, the region attribute information may be understood as, but is not limited to, basic attribute information of the image region, such as length, width, area, image depth, image color, element distribution information, and the like.
[0088] It should be noted that the regional center point of each target image area is obtained; the target offset between each regional center point and the wheel contact point of the corresponding target vehicle is calculated using the target image substructure; and the position information of the wheel contact point of each target vehicle is obtained based on each target offset and the regional attribute information of the corresponding target image area.
[0089] To further illustrate, the optional Figure 4 The scenario shown, continuing with e.g. Figure 6 As shown, the area center point 602 of the target image area 408 is obtained; the target offset between the area center point 602 and the corresponding wheel contact center point 406 of the target vehicle 404 is calculated using the target image substructure; and the position information of the wheel contact point of the target vehicle 404 is obtained based on the target offset and the area attribute information of the corresponding target image area 408.
[0090] Through the embodiments provided in the present application, the regional center point of each target image area is obtained; the target offset between each regional center point and the wheel contact point of the corresponding target vehicle is calculated using the target image substructure; and the position information of the wheel contact point of each target vehicle is obtained based on each target offset and the regional attribute information of the corresponding target image area, thereby achieving the effect of improving the accuracy of the position information of the wheel contact point.
[0091] As an optional solution, in order to improve the efficiency of obtaining information on the wheel contact points, the above-mentioned method for obtaining the wheel contact points can also be completed by means of a model, but is not limited to it. For example, the candidate recognition image is input into an image recognition model, wherein the image recognition model is a neural network model obtained after training using multiple sample image data and used to identify the position information of the wheel contact points in the image; and then the recognition result output by the image recognition model is obtained, wherein the recognition result includes the position information of the wheel contact points of each target vehicle in the candidate recognition image.
[0092] As an optional solution, before inputting the candidate recognition image into the image recognition model, include:
[0093] S1, obtaining a plurality of sample image data;
[0094] S2, marking the image data representing the wheel contact point in each sample image data to obtain a plurality of marked sample image data, wherein each marked sample image data includes a marked wheel contact point identifier;
[0095] S3, inputting the labeled multiple sample image data into the initial image recognition model to train the image recognition model.
[0096] Optionally, in this embodiment, the marked wheel contact point identification is used as a sample input to the initial image recognition model to train an image recognition model that can recognize the wheel contact point.
[0097] As an optional solution, multiple labeled sample image data are input into the initial image recognition model to train the image recognition model, including:
[0098] Repeat the following steps until you get the image recognition model:
[0099] S1, determining current sample image data from a plurality of marked sample image data and determining a current image recognition model, wherein the current sample image data includes a marked current wheel contact point identifier;
[0100] S2, identifying first result data representing the center point of the wheel contacting the ground in the current sample image data using the current image recognition model;
[0101] S3, processing the first result data using the current image recognition model to obtain second result data representing position information of the wheel contact point;
[0102] S4, if the second result data does not meet the recognition convergence condition, obtaining the next sample image data as the current sample image data;
[0103] S5. When the second result data reaches a recognition convergence condition, determine that the current image recognition model is the image recognition model.
[0104] As an optional solution, before inputting the candidate recognition image into the image recognition model, include:
[0105] S1, obtaining a plurality of sample image data;
[0106] S2, marking the first image data for indicating the wheel contact point and the second image data for indicating the wheel contact center point in each sample image data to obtain a plurality of marked sample image data, wherein each marked sample image data includes a marked wheel contact point identifier and a marked wheel contact center identifier;
[0107] S3, inputting the labeled multiple sample image data into the initial image recognition model to train the image recognition model.
[0108] Optionally, in this embodiment, the marked wheel contact point identifier and the marked wheel contact center identifier are used as samples to input into the initial image recognition model to train an image recognition model that can identify the wheel contact point and the wheel contact center.
[0109] As an optional solution, multiple labeled sample image data are input into the initial image recognition model to train the image recognition model, including:
[0110] Repeat the following steps until you get the image recognition model:
[0111] S1, determining current sample image data from a plurality of marked sample image data and determining a current image recognition model, wherein the current sample image data includes a marked current wheel contact point identifier;
[0112] S2, identifying first result data representing the center point of the wheel contacting the ground in the current sample image data using the current image recognition model;
[0113] S3, if the first result data does not meet the second convergence condition, obtaining the next sample image data as the current sample image data;
[0114] S4, when the first result data meets the second convergence condition, processing the first result data using the current image recognition model to obtain second result data for indicating position information of the wheel contact point;
[0115] S5, if the second result data does not meet the second convergence condition, obtaining the next sample image data as the current sample image data;
[0116] S6. When the second result data reaches a second convergence condition, determine that the current image recognition model is the image recognition model.
[0117] As an optional solution, obtaining the position information of the wheel contact point of each target vehicle using the target image substructure includes at least one of the following:
[0118] S1, using the target image substructure to obtain first position information of the visible wheel contact point of each target vehicle;
[0119] S2, using the target image substructure to obtain second position information of the invisible wheel contact point of each target vehicle.
[0120] Optionally, in this embodiment, in order to improve the comprehensiveness of the wheel contact point information, the wheel contact point information may, but is not limited to, carry the wheel contact point and specific attribute data of the wheel contact point, such as the left front wheel, left rear wheel, right front wheel and right rear wheel, or the visible wheel contact point, invisible wheel contact point, etc.
[0121] It should be noted that the target image substructure is used to obtain the first position information of the visible wheel contact point of each target vehicle; and the target image substructure is used to obtain the second position information of the invisible wheel contact point of each target vehicle.
[0122] To illustrate further, the optional Figure 3 As shown, the solid-line circles of the wheel contact point 310 - 1 and the wheel contact point 310 - 2 represent visible wheel contact points, and the dotted-line circles represent invisible wheel contact points.
[0123] Through the embodiments provided in the present application, the first position information of the visible wheel contact point of each target vehicle is obtained by using the target image substructure; the second position information of the invisible wheel contact point of each target vehicle is obtained by using the target image substructure, thereby achieving the effect of improving the comprehensiveness of the wheel contact point information.
[0124] As an optional solution, using the target image substructure to obtain position information of the wheel contact points of each target vehicle includes: using the target image substructure to obtain a contact point set corresponding to each target vehicle, wherein the contact point set includes a plurality of wheel contact points;
[0125] As an optional solution, after obtaining the position information of the wheel contact points of each target vehicle using the target image substructure, the method includes: determining a target contact point set from the contact point sets corresponding to each target vehicle, wherein the number of wheel contact points in the target contact point set is greater than a second threshold; and filtering the wheel contact points in the target contact point set using a non-maximum suppression method, wherein the number of wheel contact points in the filtered target contact point set is less than or equal to the second threshold.
[0126] Optionally, in this embodiment, to improve the accuracy of the wheel contact point information, it is also possible, but not limited to, after obtaining the position information of the wheel contact point of each target vehicle using the target image substructure, to suppress some contact point instances with high similarity through non-maximum suppression, and what remains is the detected contact point instance.
[0127] Through the embodiments provided in the present application, a target image substructure is used to obtain a set of contact points corresponding to each target vehicle, wherein the contact point set includes multiple wheel contact points; a target contact point set is determined from the set of contact points corresponding to each target vehicle, wherein the number of wheel contact points in the target contact point set is greater than a second threshold; and the wheel contact points in the target contact point set are filtered using a non-maximum suppression method, wherein the number of wheel contact points in the filtered target contact point set is less than or equal to the second threshold, thereby achieving the effect of improving the accuracy of the wheel contact point information.
[0128] As an optional solution, for ease of understanding, the above-mentioned wheel contact point acquisition method is applied to a real-time end-to-end wheel contact point detection scenario. For example, a detection model (image substructure) is constructed based on an ultra-lightweight convolutional neural network, and a key point (wheel contact center point)-based method is used to detect the wheel contact point. This supports the prediction of occluded contact points. In addition, due to the use of a structural selection method based on an anchor point mechanism to determine the detection model (target image substructure) corresponding to each target vehicle, the processing efficiency of the detection model is improved. For example, a target vehicle only needs to be processed using part of the detection submodel in the overall detection model. In this way, deployment can be completed on CPU devices with low computing power (such as vehicle computers), and further real-time detection of the wheel contact points of other vehicles in the current driving environment can be carried out.
[0129] Optionally, in this embodiment, keypoint annotation is performed, distinguishing between the left front wheel, left rear wheel, right front wheel, and right rear wheel. Visible wheel contact points can be directly annotated, while invisible contact points are annotated based on experience and accompanied by a visibility attribute. Annotations are visualized, such as circles of different colors representing the marked contact points, while invisible points are displayed with a white outer circle.
[0130] Optionally, in this embodiment, a detection network is built based on depthwise separable convolution, and then based on a feature pyramid structure (different from the original feature pyramid solution, the feature pyramid here also uses depthwise separable convolution instead of ordinary convolution to build, which has the advantage of low computational complexity), the last layer of feature maps of the network is used as the detection head to predict the touchdown instance; the detailed configuration of the network is shown in the following table (1), where the linear bottleneck layer is the basic structure in mobilenetV2, and the subsequent expansion coefficient, number of convolution kernels, number of repetitions and stride are the parameters of the linear bottleneck layer. The feature pyramid is √, which means that the feature fusion needs to go through the feature pyramid. For further example, the optional example Figure 7 The overall network diagram shown in the figure includes the backbone and feature pyramid structures, where the upsampling structure of the feature pyramid is based on the concatenation of depth-separable upsampling and bilinear upsampling, and Figure 7 The right subfigure in the figure illustrates the overall training process for the aforementioned wheel contact point acquisition method. First, the original image undergoes preprocessing and is fed into the detection network for forward computation, generating outputs from different detection heads. Next, based on the network outputs and pre-labeled labels, a loss is constructed to calculate gradients and update the network. The overall computational overhead of this solution is 52 MFLOPs, a very low floating-point computational level sufficient for real-time execution on low-end CPUs.
[0131] Input size operate Expansion Factor Number of convolution kernels Number of repetitions stride length Feature Pyramid 144×256×3 3x3 convolution - 16 1 2 - 72×128×16 Linear bottleneck layer 1 8 1 1 - 72×128×8 Linear bottleneck layer 6 8 2 2 - 36×64×16 Linear bottleneck layer 6 16 3 2 -
[0132] Table (1)
[0133] Input size operate Expansion Factor Number of convolution kernels Number of repetitions stride length Feature Pyramid 18×32×24 Linear bottleneck layer 6 24 4 2 √ 9×16×14 Linear bottleneck layer 6 32 3 1 - 9×16×14 Linear bottleneck layer 6 56 2 1 √ Continued Table (1)
[0134] Optionally, in this embodiment, the key point regression prediction based on the anchor point mechanism is as follows: Figure 8 As shown, the feature map of fixed height and width is regarded as a grid, which is evenly distributed on the original image 802 (such as the image area enclosed by the white grid lines). To further illustrate, it is necessary to first determine where the touchdown points are located and where they are not. For example, if the centers of the four touchdown points of a car in the original image 802 are located on a certain grid, the feature vector represented by that grid is responsible for predicting the four touchdown points of the car. Alternatively, it can be understood that the feature vector at that grid is used to predict the four touchdown points of the car.
[0135] Furthermore, for the aforementioned network, this embodiment generates feature maps for two detection heads, with dimensions of 9×16 and 18×32, respectively. The center of a specific vehicle contact point instance will fall into both layers of the detection head feature maps. This can lead to ambiguity as to which detection head detected the vehicle's wheel contact point. Based on this, an anchor box mechanism is introduced. These anchor boxes are hypothetically distributed within each grid in the detection head feature map (with the center of the anchor box aligned with the grid center). For a 9×16 feature map (relative to the network input), a 300×300 width and height anchor box (relative to 1280×720) is predefined. For an 18×32 feature map, an 80×80 width and height anchor box is predefined. The largest IOU between the largest bounding rectangle formed by the four wheel contact point instances and the predefined anchor boxes of different detection heads is used to predict the vehicle's contact point. This regulation allows the dense detection heads (18×32) to predict the touchdown point of small vehicles, and the sparser detection heads (9×16) to predict the touchdown point of large vehicles.
[0136] Optionally, in this embodiment, regarding how the network predicts whether a certain point in the feature map contains the center of the wheel contact point of a vehicle instance, a binary classification based on cross entropy is used for prediction, as shown in the following formula (1):
[0137] (1)
[0138] in, The confidence score of whether the center of the touchdown point exists here is predicted by the network. Index 0 represents the background, and index 1 represents the center of the touchdown point. is the one-hot true value label of the touchdown point. If there is no touchdown point center, it is [1, 0]. If there is a touchdown point, it is [0, 1]. Represents the cross entropy calculation function.
[0139] Optionally, in this embodiment, for the feature vector responsible for predicting the vehicle touchdown point, it is necessary to predict the locations of the four touchdown points and whether they are visible. For the location of the point, a method based on the offset of the key point is used to make the prediction. For example, the offset based on the center of the grid is used to represent its location. In order to balance the dimensions of the offset calculation of the parking space key point by different detection heads, the offset is relative to the height and width of the anchor point to avoid the dimension problem that the calculated value of the touchdown point of a small vehicle is smaller than that of the calculated value of the touchdown point of a large vehicle, such as Figure 8 As shown, the touchdown point of the vehicle instance is predicted by the feature vector represented by the target grid 804. Taking the right rear wheel point of the vehicle instance as an example, its coordinates are represented by x_offset and y_offset, and the corresponding loss is shown in the following formula (2):
[0140] (2)
[0141] in, is the width and height of the anchor box, is the actual pixel coordinate of the touchdown point, is the pixel coordinate of the center of the grid responsible for predicting the touchdown point, is the coordinate offset of the touchdown point predicted by the network, is the true value of the coordinate offset, which is obtained based on the actual pixel coordinates of the touchdown point and the pixel coordinates of the grid center.
[0142] Optionally, in this embodiment, the attribute of whether the wheel contact point is visible can be predicted using binary cross entropy, as shown in the following formula (3):
[0143] (3)
[0144] in, is the network's predicted confidence score for whether this point is visible, index 0 represents invisible, index 1 represents visible, is the one-hot true value label, if the point is invisible, it is [1, 0], if it is visible, it is [0, 1].
[0145] Optionally, during the detection phase of this embodiment, the captured vehicle-side image may be fed into a network, which then outputs a two-layer feature map for the detection head. Next, each feature vector (grid) in the feature map is traversed sequentially. If the confidence level of the touchdown point center in a feature vector exceeds a preset threshold, the feature vector is considered to represent four wheel touchdown points. The specific locations of the four touchdown points are then determined by obtaining the coordinate offsets of the four touchdown points. After traversing the two-layer feature map, a series of touchdown point instances are obtained (four groups, including left front, right front, left rear, and right rear touchdown points). Non-maximum suppression is then used to suppress those touchdown point instances with high similarity. The remaining touchdown point instances are the detected touchdown point instances. The non-maximum suppression method is essentially the same as that used in object detection, except that the calculation in object detection is based on a rectangular box, while this method is based on a closed graph formed by the touchdown point instances.
[0146] Through the embodiments provided by the present application, the wheel contact point is detected end-to-end in units of vehicle instances; furthermore, the algorithm of the above-mentioned method for obtaining the wheel contact point can be calculated with the help of only one forward transmission calculation of the network; it can be deployed on a CPU device to achieve real-time performance; in addition, the overall computational complexity of the above-mentioned method for obtaining the wheel contact point is low, and it can be run on devices with lower computing power, and most detection functions can be met in a single-core operation mode; further, the accuracy and recall rate of the above-mentioned method for obtaining the wheel contact point can be maintained at a considerable level while taking into account the computational complexity, and basically meet the user's usage requirements. The accuracy and recall rate can be maintained at a considerable level while taking into account the computational complexity, and basically meet the user's usage requirements.
[0147] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0148] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0149] According to another aspect of the embodiment of the present application, there is also provided a wheel contact point acquisition device for implementing the above-mentioned wheel contact point acquisition method. Figure 9 As shown, the device includes:
[0150] A first determining unit 902 is configured to, upon acquiring a candidate recognition image and identifying at least one target vehicle from the candidate recognition image, determine a wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point represents position information of centers of at least two wheel contact points of the target vehicle in the candidate recognition image;
[0151] A first acquisition unit 904 is configured to acquire at least two image substructures corresponding to the candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image region in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to obtain position information of the wheel contact point;
[0152] A second determining unit 906 is configured to determine, based on the image region where the wheel contact center is located, a target image substructure corresponding to each target vehicle from the at least two image substructures;
[0153] The second acquiring unit 908 is configured to acquire the position information of the wheel contact point of each target vehicle by using the target image substructure.
[0154] Optionally, in this embodiment, the above-mentioned wheel contact point acquisition device can be used, but is not limited to, in a real-time end-to-end wheel contact point detection scenario. For example, a detection model (image substructure) is constructed based on an ultra-lightweight convolutional neural network, and a device based on a key point (wheel contact center point) is used to detect the wheel contact point, supporting the prediction of occluded contact points. Moreover, since a structural selection method based on an anchor point mechanism is adopted to determine the detection model (target image substructure) corresponding to each target vehicle, the processing efficiency of the detection model is improved. For example, a target vehicle only needs to be processed using part of the detection submodel in the overall detection model. In this way, deployment can be completed on a CPU device with low computing power (such as a vehicle computer), and further real-time detection of the wheel contact points of other vehicles in the current driving environment is achieved.
[0155] Optionally, in this embodiment, the wheel contact point can, but is not limited to, help the autonomous vehicle obtain a more accurate relative position, direction, and size of the vehicle, which is of great significance to improving safety during driving; for example, after using the target image substructure to obtain the position information of the wheel contact point of each target vehicle, the position information of the wheel contact point is used to determine the spatial information of the target vehicle, wherein the spatial information can, but is not limited to, include at least one of the following: vehicle relative position, vehicle relative direction, vehicle size, vehicle driving trajectory, etc.
[0156] Optionally, in this embodiment, the above-mentioned wheel contact point acquisition device can also be used, but is not limited to, in scenarios where relatively accurate relative distance detection of other vehicles is required. For example, based on the acquired contact point, its corresponding center (pixel position) is calculated, and then converted to the world coordinate position in the vehicle body coordinate system through pre-calibrated camera parameters and ground assumptions. This can help ADAS warnings, lane-level vehicle rendering, etc. obtain accurate position information of other vehicles.
[0157] Optionally, in this embodiment, the above-mentioned wheel contact point acquisition device can also be used, but is not limited to, in scenarios where a monocular camera performs 3D target detection of a vehicle. For example, based on the acquired contact point, combined with 2D target detection and camera calibration parameters, 3D target detection information of the vehicle can be generated, including the center position (x, y, z), length, width, height, and heading angle.
[0158] Optionally, in this embodiment, the method of acquiring the candidate recognition image may be, but is not limited to, a method of acquiring images of the vehicle during operation using an on-board image acquisition device, or a method of acquiring road condition images using a roadside camera.
[0159] Optionally, in this embodiment, image features can be understood as, but not limited to, features presented by image pixels in the candidate recognition image, or can be understood as features identified by image recognition technology, such as color features, shape features, local image features, image depth features, edge features, linear features, texture features, etc.; determining the wheel contact center point of each target vehicle based on the image features of the candidate recognition image can be accomplished by, but not limited to, using an image recognition model, for example, first determining an initial image recognition model, and then inputting labeled sample image data for training to obtain a trained image recognition model, wherein the labeled sample image data includes image data of the target vehicle with labeled key points, and the key points include wheel contact points, and specific attribute data of the wheel contact points, such as the left front wheel, left rear wheel, right front wheel and right rear wheel, or visible wheel contact points, invisible wheel contact points, etc.
[0160] Optionally, in this embodiment, the wheel contact center point can be understood as, but not limited to, the center of at least two wheel contact points of the target vehicle. Since the at least two wheel contact points may include a currently invisible wheel contact point, in this case, it is possible but not limited to first predicting the position information of the invisible wheel contact points, and then combining the position information of the visible wheel contact points to predict the position information of the wheel contact center point.
[0161] Optionally, in this embodiment, the target image model may include, but is not limited to, at least two image substructures, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, or it can be understood that each image substructure corresponds to at least one image area in the candidate recognition image, and thus the corresponding image substructure can be directly determined when the image area is determined first, and vice versa; in addition, the image substructure is used to identify the image content information of the candidate recognition image to obtain the position information of the wheel contact point, or it can be understood that the image substructure corresponds to the image area, but the image substructure is used to identify the overall image content information of the candidate recognition image, or is not limited to the image content information of the image area corresponding to the image substructure, but can be, but is not limited to, a stronger correlation with the image content information of the corresponding image area, or the recognition accuracy of the image content information of the corresponding image area is higher, thereby improving the recognition accuracy of the wheel contact point.
[0162] Optionally, in this embodiment, based on the image region where the wheel contact center is located, a target image substructure corresponding to each target vehicle is determined from at least two image substructures. Alternatively, this can be understood as determining the target image region where the wheel contact center is located from multiple image regions, and then using the correspondence between the image region and the image substructures, determining the target image substructure corresponding to the target image region from at least two image substructures. Furthermore, since the target image region is determined by the wheel contact center, and the wheel contact center clearly belongs to a target vehicle, when the target image substructure is determined, the target vehicle corresponding to the target image substructure can also be obtained. Alternatively, this can be understood as a correspondence between the target image substructure and the target vehicle.
[0163] Optionally, in this embodiment, the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle. For example, the target image area in the candidate recognition image is first determined by using the correspondence between the target image substructure and the image area, and then the area center point of the target image area is determined; further, the anchor point mechanism is used to calculate the position offset between the wheel contact point of each target vehicle and the area center point using the target image substructure, and then the actual position information of the wheel contact point of each target vehicle is calculated based on the position offset.
[0164] It should be noted that the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle, and the anchor point mechanism is used to determine the target image substructure corresponding to each vehicle through the wheel contact center point. The target image substructure is then used to perform more targeted wheel contact point identification on the corresponding vehicle, and the position information of the identified wheel contact point is also associated with the corresponding vehicle, making the position information of the wheel contact point more comprehensive. In addition, the wheel contact center point is used to determine the target image substructure corresponding to each target vehicle from multiple image substructures, thereby improving the application efficiency of the image substructure. In this way, the above-mentioned wheel contact point acquisition device can be implemented in application scenarios with lower computing power, such as application scenarios of vehicle-mounted equipment.
[0165] For a specific embodiment, reference may be made to the example shown in the above-mentioned device for obtaining the wheel contact point, which will not be described in detail in this example.
[0166] Through the embodiments provided by the present application, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, the wheel contact center point of each target vehicle is determined based on the image features of the candidate recognition image, wherein the wheel contact center point is used to represent the position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image; at least two image substructures corresponding to the candidate recognition image are acquired, wherein each of the at least two image substructures corresponds to at least one image area in the candidate recognition image, and the image substructure is used to identify the image content information of the candidate recognition image to obtain the position information of the wheel contact point; and each target vehicle is determined from the at least two image substructures based on the image area where the wheel contact center point is located. The target image substructure corresponding to each target vehicle is used; the target image substructure is used to obtain the position information of the wheel contact point of each target vehicle; the embodiment of the present application uses the target image substructure to obtain the position information of the wheel contact point of each target vehicle, uses the anchor point mechanism, and determines the target image substructure corresponding to each vehicle through the wheel contact center point, and then uses the target image substructure to perform more targeted wheel contact point identification for each corresponding vehicle, and the position information of the identified wheel contact point is also associated with the corresponding vehicle, making the position information of the wheel contact point more comprehensive; in addition, the wheel contact center point is used to determine the target image substructure corresponding to each target vehicle from multiple image substructures, thereby improving the application efficiency of the image substructure.
[0167] As an optional solution, the second determining unit 906 includes:
[0168] A first acquisition module is configured to acquire an image region set corresponding to at least two image substructures, wherein the image region set includes at least two image regions in the candidate recognition image;
[0169] A second acquisition module is configured to acquire a target image region where the wheel contact center point corresponding to each target vehicle is located, wherein at least two image regions include the target image region;
[0170] The first determining module is configured to determine a target image substructure corresponding to each target image region from at least two image substructures.
[0171] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0172] As an optional solution, the first acquisition module includes:
[0173] A first acquisition submodule is configured to acquire a first region set corresponding to at least two first image substructures, wherein the first region set includes at least two first image regions in the candidate recognition image, and the first image regions correspond to the first image substructures;
[0174] The second acquisition submodule is used to acquire a second area set corresponding to at least two second image substructures, wherein the second area set includes at least two second image areas in the candidate recognition image, the second image areas correspond to the second image substructures, and the area range of the second image areas is larger than the area range of the first image areas.
[0175] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0176] As an optional solution, the first determination module includes:
[0177] A first determining submodule is configured to determine a target first substructure and a target second substructure corresponding to each target image region, wherein the target image substructure includes the target first substructure and the target second substructure; or
[0178] The second determination submodule is used to determine the target first substructure corresponding to each target image area when the space occupied by the target vehicle is less than the first threshold, and to determine the target second substructure corresponding to each target image area when the space occupied by the target vehicle is greater than or equal to the first threshold.
[0179] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0180] As an optional solution, the second obtaining unit 908 includes:
[0181] A third acquisition module is used to obtain the center point of each target image area;
[0182] a calculation module, configured to calculate a target offset between a center point of each region and a wheel contact point of a corresponding target vehicle using the target image substructure;
[0183] The fourth acquisition module is used to acquire the position information of the wheel contact point of each target vehicle according to each target offset and the area attribute information of the corresponding target image area.
[0184] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0185] As an optional solution, the device further includes:
[0186] An input unit, configured to input a candidate recognition image into an image recognition model, wherein the image recognition model is a neural network model trained using a plurality of sample image data and configured to recognize position information of a wheel contact point in an image;
[0187] The third acquisition unit is used to acquire the recognition result output by the image recognition model, wherein the recognition result includes the position information of the wheel contact point of each target vehicle in the candidate recognition image.
[0188] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0189] As an optional solution, it includes:
[0190] a fourth acquisition unit, configured to acquire a plurality of sample image data before inputting the candidate recognition image into the image recognition model;
[0191] a first labeling unit, configured to label image data representing a wheel contact point in each sample image data before inputting the candidate recognition image into the image recognition model, to obtain a plurality of labeled sample image data, wherein each labeled sample image data includes a labeled wheel contact point identifier;
[0192] The first training unit is used to input the labeled multiple sample image data into the initial image recognition model before inputting the candidate recognition image into the image recognition model, so as to train the image recognition model.
[0193] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0194] As an optional solution, the first training unit includes:
[0195] The first repetition module is used to repeatedly perform the following steps until an image recognition model is obtained:
[0196] a second determining module, configured to determine current sample image data from the labeled plurality of sample image data and determine a current image recognition model, wherein the current sample image data includes a labeled current wheel contact point identifier;
[0197] A first recognition module is configured to recognize, by using a current image recognition model, first result data representing a wheel contact center point in the current sample image data;
[0198] a first processing module, configured to process the first result data using a current image recognition model to obtain second result data representing position information of a wheel contact point;
[0199] A second processing module is configured to obtain the next sample image data as the current sample image data if the second result data does not meet the recognition convergence condition;
[0200] The third processing module is configured to determine that the current image recognition model is the image recognition model when the second result data reaches a recognition convergence condition.
[0201] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0202] As an optional solution, it includes:
[0203] a fifth acquisition unit, configured to acquire a plurality of sample image data before inputting the candidate recognition image into the image recognition model;
[0204] a second labeling unit, configured to label first image data representing a wheel contact point and second image data representing a wheel contact center point in each sample image data before inputting the candidate recognition image into the image recognition model, to obtain a plurality of labeled sample image data, wherein each labeled sample image data includes a labeled wheel contact point identifier and a labeled wheel contact center identifier;
[0205] The second training unit is used to input the labeled multiple sample image data into the initial image recognition model before inputting the candidate recognition image into the image recognition model, so as to train the image recognition model.
[0206] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0207] As an optional option, the second training unit includes:
[0208] The second repetition module is used to repeatedly perform the following steps until an image recognition model is obtained:
[0209] a third determining module, configured to determine current sample image data from the labeled plurality of sample image data and determine a current image recognition model, wherein the current sample image data includes a labeled current wheel contact point identifier;
[0210] A second recognition module is configured to recognize, by using a current image recognition model, first result data representing a wheel contact center point in the current sample image data;
[0211] A fourth processing module, configured to obtain next sample image data as current sample image data if the first result data does not meet the second convergence condition;
[0212] a fifth processing module, configured to process the first result data using the current image recognition model to obtain second result data representing position information of a wheel contact point when the first result data meets a second convergence condition;
[0213] A sixth processing module, configured to obtain next sample image data as current sample image data if the second result data does not meet the second convergence condition;
[0214] The seventh processing module is used to determine that the current image recognition model is the image recognition model when the second result data meets the second convergence condition.
[0215] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0216] As an optional solution, the second obtaining unit 908 includes at least one of the following:
[0217] A first position module is used to obtain first position information of a wheel contact point visible to each target vehicle using the target image substructure;
[0218] The second position module is used to obtain second position information of the invisible wheel contact point of each target vehicle by using the target image substructure.
[0219] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0220] As an optional solution, the second acquisition unit 908 includes: a fifth acquisition module, configured to acquire a contact point set corresponding to each target vehicle using the target image substructure, wherein the contact point set includes a plurality of wheel contact points;
[0221] The device includes: after obtaining the position information of the wheel contact points of each target vehicle using a target image substructure, determining a target contact point set from the contact point sets corresponding to each target vehicle, wherein the number of wheel contact points in the target contact point set is greater than a second threshold; and filtering the wheel contact points in the target contact point set using a non-maximum suppression method, wherein the number of wheel contact points in the filtered target contact point set is less than or equal to the second threshold.
[0222] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for obtaining the wheel contact point, which will not be described in detail in this example.
[0223] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned method for obtaining the wheel contact point is also provided. Figure 10As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0224] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0225] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0226] S1, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, determining the wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point is used to represent position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image;
[0227] S2, obtaining at least two image substructures corresponding to the candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image region in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to obtain position information of the wheel contact point;
[0228] S3, determining a target image substructure corresponding to each target vehicle from the at least two image substructures according to the image region where the wheel contact center point is located;
[0229] S4, using the target image substructure to obtain the position information of the wheel contact point of each target vehicle.
[0230] Alternatively, those skilled in the art will appreciate that Figure 10 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 10 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 10 More or fewer components (such as network interfaces, etc.) as shown in the Figure 10 Different configurations shown.
[0231] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for obtaining the wheel contact point in the embodiment of the present application. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, realizing the above-mentioned method for obtaining the wheel contact point. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include a memory remotely located relative to the processor 1004, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1002 can be used specifically, but not limited to, to store information such as candidate recognition images, target image substructures, and location information of the wheel contact point. As an example, if Figure 10 As shown, the memory 1002 may include, but is not limited to, the first determination unit 902, the first acquisition unit 904, the second determination unit 906, and the second acquisition unit 908 of the device for acquiring the wheel contact point. Furthermore, the memory 1002 may also include, but is not limited to, other modules and units of the device for acquiring the wheel contact point, which will not be further described in this example.
[0232] Optionally, the transmission device 1006 is configured to receive or transmit data via a network. Specific examples of the network include wired networks and wireless networks. In one embodiment, the transmission device 1006 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to enable communication with the Internet or a local area network. In another embodiment, the transmission device 1006 is a radio frequency (RF) module configured to communicate with the Internet wirelessly.
[0233] In addition, the electronic device further includes: a display 1008 for displaying information such as the candidate recognition image, target image substructure, and wheel contact point location information; and a connection bus 1010 for connecting various module components in the electronic device.
[0234] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, which may be a distributed system formed by connecting multiple nodes via network communications. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0235] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instructions containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions provided in the embodiments of the present application are performed.
[0236] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0237] It should be noted that the computer system of the electronic device is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0238] A computer system includes a central processing unit (CPU), which performs various actions and processes based on programs stored in read-only memory (ROM) or loaded from storage into random access memory (RAM). RAM also stores various programs and data required for system operation. The CPU, ROM, and RAM are connected to each other via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0239] The following components are connected to the input / output interface: an input section including a keyboard and mouse; an output section including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; a storage section including a hard disk; and a communication section including network interface cards such as local area network cards and modems. The communication section performs communication processing via a network such as the Internet. Drives are also connected to the input / output interface as needed. Removable media such as magnetic disks, optical disks, magneto-optical disks, and semiconductor memories are installed in the drive as needed, allowing computer programs read from these media to be installed in the storage section as needed.
[0240] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication portion, and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions defined in the system of the present application are performed.
[0241] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementations described above.
[0242] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0243] S1, when a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, determining the wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point is used to represent position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image;
[0244] S2, obtaining at least two image substructures corresponding to the candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image region in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to obtain position information of the wheel contact point;
[0245] S3, determining a target image substructure corresponding to each target vehicle from the at least two image substructures according to the image region where the wheel contact center point is located;
[0246] S4, using the target image substructure to obtain the position information of the wheel contact point of each target vehicle.
[0247] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0248] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0249] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.
[0250] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0251] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0252] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0253] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0254] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for obtaining a wheel contact point, characterized in that: include: When a candidate recognition image is acquired and at least one target vehicle is identified from the candidate recognition image, a first image recognition model is used to determine the wheel contact center point of each target vehicle based on image features of the candidate recognition image, wherein the wheel contact center point represents the position information of the center of at least two wheel contact points of the target vehicle in the candidate recognition image, and the first image recognition model is a neural network model obtained by training using multiple sample image data and is used to recognize the position information of the wheel contact center point in the image; Acquiring at least two image substructures corresponding to a candidate recognition image, wherein each of the at least two image substructures corresponds to at least one image region in the candidate recognition image, and the image substructures are used to identify image content information of the candidate recognition image to obtain position information of the wheel contact point; Determining, according to the image region where the wheel contact center point is located, a target image substructure corresponding to each target vehicle from the at least two image substructures; The target image substructure is used to obtain position information of the wheel contact point of each target vehicle.
2. The method according to claim 1, characterized in that The determining, based on the image area where the wheel contact center point is located, a target image substructure corresponding to each target vehicle from the at least two image substructures includes: Acquire an image region set corresponding to the at least two image substructures, wherein the image region set includes at least two image regions in the candidate recognition image; Acquire a target image region where the wheel contact center point corresponding to each target vehicle is located, wherein the at least two image regions include the target image region; The target image substructure corresponding to each target image region is determined from the at least two image substructures.
3. The method according to claim 2, characterized in that The acquiring of the image region set corresponding to the at least two image substructures includes: Acquire a first region set corresponding to at least two first image substructures, wherein the first region set includes at least two first image regions in the candidate recognition image, and the first image regions correspond to the first image substructures; Obtain a second region set corresponding to at least two second image substructures, wherein the second region set includes at least two second image regions in the candidate recognition image, the second image regions correspond to the second image substructures, and the area range of the second image regions is larger than the area range of the first image region.
4. The method according to claim 3, characterized in that The determining, from the at least two image substructures, the target image substructure corresponding to each target image region includes: Determine a target first substructure and a target second substructure corresponding to each target image region, wherein the target image substructure includes the target first substructure and the target second substructure; or When the space occupied by the target vehicle is less than a first threshold, the target first substructure corresponding to each target image area is determined; and when the space occupied by the target vehicle is greater than or equal to the first threshold, the target second substructure corresponding to each target image area is determined.
5. The method according to claim 2, characterized in that The acquiring the position information of the wheel contact point of each target vehicle by using the target image substructure includes: Obtaining a center point of each target image area; Calculating a target offset between each of the region center points and the wheel contact point of the corresponding target vehicle using the target image substructure; The position information of the wheel contact point of each target vehicle is obtained according to each target offset and the area attribute information of the corresponding target image area.
6. The method according to claim 1, wherein The method further comprises: Inputting the candidate recognition image into a second image recognition model, wherein the image recognition model is a neural network model obtained by training with a plurality of sample image data and used to recognize the position information of the wheel contact point in the image; Acquire a recognition result output by the second image recognition model, wherein the recognition result includes position information of the wheel contact point of each of the target vehicles in the candidate recognition image.
7. The method according to claim 6, characterized in that Before inputting the candidate recognition image into the image recognition model, the method includes: acquiring the plurality of sample image data; Marking the image data used to represent the wheel contact point in each of the sample image data to obtain the plurality of marked sample image data, wherein each marked sample image data includes a marked wheel contact point identifier; The labeled plurality of sample image data are input into an initial image recognition model to train and obtain the second image recognition model.
8. The method according to claim 7, characterized in that Inputting the labeled plurality of sample image data into an initial image recognition model to train the second image recognition model includes: Repeat the following steps until the second image recognition model is obtained: Determining current sample image data from the labeled plurality of sample image data, and determining a current image recognition model, wherein the current sample image data includes a labeled current wheel contact point identifier; Identifying, by the current image recognition model, first result data representing the wheel contact center point in the current sample image data; Processing the first result data using the current image recognition model to obtain second result data representing position information of the wheel contact point; If the second result data does not meet the recognition convergence condition, obtaining the next sample image data as the current sample image data; When the second result data meets the recognition convergence condition, the current image recognition model is determined to be the second image recognition model.
9. The method according to claim 6, characterized in that Before inputting the candidate recognition image into the second image recognition model, the method includes: acquiring the plurality of sample image data; Marking the first image data for representing the wheel contact point and the second image data for representing the wheel contact center point in each of the sample image data to obtain the plurality of marked sample image data, wherein each marked sample image data includes a marked wheel contact point identifier and a marked wheel contact center identifier; The labeled plurality of sample image data are input into an initial image recognition model to train and obtain the second image recognition model.
10. The method according to claim 9, characterized in that Inputting the labeled plurality of sample image data into an initial image recognition model to train the second image recognition model includes: Repeat the following steps until the second image recognition model is obtained: Determining current sample image data from the labeled plurality of sample image data, and determining a current image recognition model, wherein the current sample image data includes a labeled current wheel contact point identifier; Identifying, by the current image recognition model, first result data representing the wheel contact center point in the current sample image data; When the first result data does not meet the second convergence condition, obtaining next sample image data as the current sample image data; When the first result data meets the second convergence condition, the first result data is processed by the current image recognition model to obtain second result data representing position information of the wheel contact point; If the second result data does not meet the second convergence condition, obtaining next sample image data as the current sample image data; When the second result data meets the second convergence condition, the current image recognition model is determined to be the second image recognition model.
11. The method according to any one of claims 1 to 10, characterized in that The acquiring the position information of the wheel contact point of each target vehicle by using the target image substructure includes at least one of the following: Acquire first position information of the wheel contact point visible to each target vehicle using the target image substructure; The target image substructure is used to obtain second position information of the invisible wheel contact point of each target vehicle.
12. The method according to any one of claims 1 to 10, characterized in that The acquiring the position information of the wheel contact point of each target vehicle by using the target image substructure includes: acquiring a contact point set corresponding to each target vehicle by using the target image substructure, wherein the contact point set includes a plurality of wheel contact points; After obtaining the position information of the wheel contact points of each target vehicle using the target image substructure, the method further includes: determining a target contact point set from the contact point sets corresponding to each target vehicle, wherein the number of the wheel contact points in the target contact point set is greater than a second threshold; and filtering the wheel contact points in the target contact point set using a non-maximum suppression method, wherein the number of the wheel contact points in the filtered target contact point set is less than or equal to the second threshold.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the method according to any one of claims 1 to 12 is executed when the program is executed.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 12 through the computer program.
Citation Information
Patent Citations
Method and device for determining parking behavior, medium and equipment
CN110176151A
Hub motor driving independent suspension mechanism of virtual rail train and design method of hub motor driving independent suspension mechanism
CN111523173A