Training method, device, electronic device and storage medium for autonomous driving model
Through multiple rounds of iterative retrieval and screening of candidate image sets, the problems of high difficulty and low accuracy in autonomous driving model training were solved, achieving more efficient training results.
Patent Information
- Application Number
- CN202211521243.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-11-30
AI Technical Summary
The training of existing autonomous driving models has the problems of high training difficulty and low training accuracy.
Through multiple rounds of iterative retrieval and screening, a set of candidate images is obtained to form a sample image set for training the autonomous driving model.
It reduces the difficulty of image retrieval, improves training accuracy, and enhances the training effect of autonomous driving models.
Smart Images

Figure CN115713749B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of autonomous driving and deep learning technology, and in particular to a training method, device, electronic device, storage medium and computer program product for an autonomous driving model. Background Art
[0002] With the continuous development of artificial intelligence (AI) technology, autonomous driving models are widely used in the automotive sector, offering advantages such as high automation and intelligence. For example, image data can be input into autonomous driving models, which then identify obstacle locations and perform route planning. However, the training of autonomous driving models in related technologies suffers from high difficulty and low accuracy. Summary of the Invention
[0003] The present disclosure provides a training method, apparatus, electronic device, storage medium, and computer program product for an autonomous driving model.
[0004] According to one aspect of the present disclosure, a training method for an autonomous driving model is provided, including: obtaining a target scene to be trained for the autonomous driving model; in the case of a first round of retrieval, retrieving a first candidate image set from a total image set based on the target scene; in the case of an Nth round of retrieval, retrieving an Nth candidate image set from the total image set based on the N-1th candidate image set retrieved in the N-1th round and the target scene, wherein 2≤N≤M, and M is an integer greater than 1; based on the M candidate image sets, obtaining a first sample image set corresponding to the target scene; and training the autonomous driving model based on the first sample image set.
[0005] According to another aspect of the present disclosure, a training device for an autonomous driving model is provided, comprising: a first acquisition module for acquiring a target scene to be trained for the autonomous driving model; a retrieval module for retrieving a first candidate image set from a total image set based on the target scene in a first round of retrieval; the retrieval module is further used to retrieve an Nth candidate image set from the total image set based on the N-1th candidate image set retrieved in the N-1th round of retrieval and the target scene, wherein 2≤N≤M, and M is an integer greater than 1; a second acquisition module for obtaining a first sample image set corresponding to the target scene based on the M candidate image sets; and a training module for training the autonomous driving model based on the first sample image set.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a training method for an autonomous driving model.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute a training method for an autonomous driving model.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of a training method for an autonomous driving model.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0011] Figure 1 is a flowchart of a method for training an autonomous driving model according to the first embodiment of the present disclosure;
[0012] Figure 2 is a schematic diagram of a training method for an autonomous driving model according to a second embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of a training method for an autonomous driving model according to a third embodiment of the present disclosure;
[0014] Figure 4 is a schematic diagram of a training method for an autonomous driving model according to a fourth embodiment of the present disclosure;
[0015] Figure 5 is a block diagram of a training device for an autonomous driving model according to a first embodiment of the present disclosure;
[0016] Figure 6 This is a block diagram of an electronic device used to implement the training method of the autonomous driving model of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] Artificial Intelligence (AI) is a discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Currently, AI technology has been widely used due to its high degree of automation, high precision, and low cost.
[0019] Autonomous driving is a complex integration of cutting-edge disciplines, including sensors, computers, artificial intelligence, communications, navigation and positioning, pattern recognition, machine vision, and intelligent control. Since the beginning of the 21st century, with the substantial increase in physical computing power, the rapid development of dynamic vision technology, and the rapid advancement of artificial intelligence, key technologies such as route navigation, obstacle avoidance, and emergency decision-making have been mastered, leading to breakthroughs in autonomous driving technology.
[0020] DL (Deep Learning) is a new research direction in the field of ML (Machine Learning). It is a science that learns the inherent laws and representation levels of sample data, enabling machines to have analytical and learning capabilities like humans and recognize data such as text, images, and sounds. It is widely used in speech and image recognition.
[0021] Figure 1 3 is a flowchart of a method for training an autonomous driving model according to the first embodiment of the present disclosure.
[0022] like Figure 1 As shown, the training method of the autonomous driving model of the first embodiment of the present disclosure includes:
[0023] S101, obtaining a target scene for the autonomous driving model to be trained.
[0024] It should be noted that the execution entity of the autonomous driving model training method of the disclosed embodiments may be a hardware device with data processing capabilities and / or the necessary software to drive the hardware device. Optionally, the execution entity may include a workstation, server, computer, user terminal, and other intelligent device. User terminals include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals.
[0025] It should be noted that there are not too many restrictions on the autonomous driving model. For example, the autonomous driving model may include an obstacle recognition model, a route planning model, etc.
[0026] It should be noted that there are not too many restrictions on the target scenarios. For example, the target scenarios include cargo scenes (such as automatic delivery vehicles, automatic sanitation vehicles), passenger scenes, cleaning scenes (such as cleaning robots), monitoring scenes (such as monitoring robots), etc.
[0027] S102 : In the case of the first round of retrieval, the first candidate image set is retrieved from the total image set based on the target scene.
[0028] It should be noted that the total image set includes a large number of images. Images can be pre-captured by cameras on multiple autonomous driving objects and added to the total image set. It should be noted that there are no specific limitations on autonomous driving objects; for example, autonomous driving objects can include vehicles, robots, etc. There are also no specific limitations on the images in the total image set; for example, images can include two-dimensional images, three-dimensional images, etc.
[0029] In one embodiment, retrieving a first candidate image set from the total image set based on the target scene includes retrieving the first candidate image set from the total image set based on a feature representation of the target scene. It should be noted that the feature representation is not particularly limited; for example, the feature representation may be a feature vector.
[0030] In some examples, a large multimodal feature extraction model can be used to extract features from a target scene to obtain a feature representation of the target scene. It should be noted that the large multimodal feature extraction model can be trained based on multimodal sample data, and there are no specific restrictions on the training process.
[0031] In some examples, based on the feature representation of the target scene, a first candidate image set is retrieved from the total image set, including obtaining a second feature representation of each image in the total image set, obtaining the similarity between the feature representation of the target scene and the second feature representation, and adding images with a similarity greater than a set threshold to the first candidate image set.
[0032] In one embodiment, retrieving the first candidate image set from the total image set based on the target scene includes randomly extracting the first first image set from the total image set, and retrieving the first candidate image set from the first first image set based on the target scene.
[0033] S103 , in the case of the Nth round of retrieval, based on the N-1th candidate image set retrieved in the N-1th round of retrieval and the target scene, retrieve the Nth candidate image set from the total image set, where 2≤N≤M, and M is an integer greater than 1.
[0034] It should be noted that there is no excessive limitation on M, for example, M=3.
[0035] Taking M=3 as an example, in the second round of retrieval, the second candidate image set is retrieved from the total image set based on the first candidate image set and the target scene retrieved in the first round. In the third round of retrieval, the third candidate image set is retrieved from the total image set based on the second candidate image set and the target scene retrieved in the second round.
[0036] In one embodiment, retrieving the Nth candidate image set from the total image set based on the N-1th candidate image set and the target scene retrieved in the N-1th round includes retrieving the Nth candidate image set from the total image set based on a feature representation of the N-1th candidate image set and a feature representation of the target scene. Thus, this method can comprehensively consider the feature representation of the N-1th candidate image set and the feature representation of the target scene to retrieve the Nth candidate image set from the total image set, thereby improving the accuracy of the Nth round of retrieval.
[0037] In some examples, based on the feature representation of the N-1th candidate image set and the feature representation of the target scene, the Nth candidate image set is retrieved from the total image set, including performing feature fusion on the feature representation of the N-1th candidate image set and the feature representation of the target scene to obtain a fused feature representation, and retrieving the Nth candidate image set from the total image set based on the fused feature representation.
[0038] S104: Obtain a first sample image set corresponding to the target scene based on the M candidate image sets.
[0039] In one embodiment, obtaining a first sample image set corresponding to the target scene based on the M candidate image sets includes adding the M candidate sample sets to the first sample image set.
[0040] In one embodiment, obtaining a first sample image set corresponding to a target scene based on M candidate image sets includes filtering the first sample image set from the M candidate image sets. Thus, the method can further filter the first sample image set from the M candidate image sets, thereby improving the accuracy of the first sample image set.
[0041] In some examples, selecting the first sample image set from the M candidate image sets includes, in response to overlapping portions between any two candidate image sets, deleting the overlapping portions from one of the two candidate image sets, retaining the overlapping portions from the other candidate image set, and adding the images from the deleted candidate image set and the images from the other candidate image set to the first sample image set. Thus, the method ensures that there are no duplicate images in the first sample image set.
[0042] S105: Training the autonomous driving model based on the first sample image set.
[0043] It should be noted that there are not too many restrictions on the training methods of autonomous driving models. For example, the training methods may include supervised training, unsupervised training, etc.
[0044] In one embodiment, training an autonomous driving model based on a first sample image set includes inputting a first sample image from the first sample image set into the autonomous driving model, having the autonomous driving model output a prediction result, and training the autonomous driving model based on the prediction result and the annotation result. It should be noted that there are no specific limitations on the annotation results; for example, the annotation results include, but are not limited to, person detection boxes, obstacle detection boxes, and routes.
[0045] In summary, according to the training method for an autonomous driving model of the embodiments of the present disclosure, a target scene to be trained for the autonomous driving model is obtained. In the first round of retrieval, the first candidate image set is retrieved from the total image set based on the target scene. In the Nth round of retrieval, the Nth candidate image set is retrieved from the total image set based on the N-1th candidate image set retrieved in the N-1th round and the target scene, where 2≤N≤M, and M is an integer greater than 1. Based on the M candidate image sets, a first sample image set corresponding to the target scene is obtained, and the autonomous driving model is trained based on the first sample image set. Thus, the Nth round of retrieval relies on the N-1th candidate image set retrieved in the N-1th round, thereby achieving multiple rounds of iterative retrieval of the total image set. Each round of retrieval can obtain one candidate image set to obtain the first sample image set. Compared to the related art, which often only performs a single retrieval in the total image set, this reduces the difficulty of searching the total image set and improves the accuracy of image retrieval, thereby reducing the difficulty of training the autonomous driving model and improving the training accuracy of the autonomous driving model.
[0046] Figure 2 3 is a flowchart of a method for training an autonomous driving model according to the second embodiment of the present disclosure.
[0047] like Figure 2 As shown, the training method of the autonomous driving model of the second embodiment of the present disclosure includes:
[0048] S201, obtaining a target scene for the autonomous driving model to be trained.
[0049] S202 : In the case of the first round of retrieval, the first candidate image set is retrieved from the total image set based on the target scene.
[0050] The relevant contents of steps S201-S202 can be found in the above embodiment and will not be repeated here.
[0051] S203 , in the case of the Nth round of retrieval, based on the N-1th candidate image set, screen out the Nth first image set from the total image set.
[0052] In one embodiment, screening out the Nth first image set from the total image set based on the N-1th candidate image set includes screening out the Nth first image set similar to the N-1th candidate image set from the total image set.
[0053] In one embodiment, based on the N-1th candidate image set, the Nth first image set is screened out from the total image set, including obtaining a feature representation of the N-1th candidate image set, obtaining a second feature representation of each image in the total image set, obtaining a similarity between the feature representation of the N-1th candidate image set and the second feature representation, and adding images with a similarity greater than a set threshold to the 1st first image set.
[0054] In one embodiment, selecting the Nth first image set from the total image set based on the N-1th candidate image set includes selecting the Nth first image set from the total image set based on acquisition parameters of the candidate images in the N-1th candidate image set. Thus, the method can consider the acquisition parameters of the candidate images in the N-1th candidate image set to select the Nth first image set from the total image set, thereby improving the accuracy of the Nth first image set.
[0055] It should be noted that there are no excessive restrictions on the acquisition parameters. For example, the acquisition parameters include the acquisition time, the number of image frames, the camera to which the image belongs, the vehicle to which the camera belongs, etc.
[0056] In one embodiment, based on the acquisition parameters of the candidate images in the N-1th candidate image set, the Nth first image set is screened out from the total image set, including the following possible implementations:
[0057] Method 1: In response to an acquisition parameter indicating that a candidate image in the N-1th candidate image set is acquired by a target camera at a target frame, a second image set acquired by the target camera at an adjacent frame of the target frame is filtered out from the total image set, and the second image set is added to the Nth first image set.
[0058] It can be understood that the image captured by the target camera in the adjacent frame of the target frame has a high similarity with the candidate image in the N-1th candidate image set.
[0059] It should be noted that the adjacent frames of the target frame may include frames that are within a set number of frames between the target frame and the target frame. There is no excessive limitation on the set number of frames, for example, it can be 3 frames.
[0060] For example, the second candidate image set includes candidate images A and B.
[0061] If the acquisition parameters of candidate image A indicate that candidate image A was acquired by target camera 1 at the 10th frame, the second image set 1 acquired by target camera 1 at the 7th frame, 8th frame, 9th frame, 11th frame, 12th frame, and 13th frame can be filtered out from the total image set, and the second image set 1 can be added to the third first image set.
[0062] If the acquisition parameters of candidate image B indicate that candidate image B was acquired by target camera 2 at the 5th frame, the second image set 2 acquired by target camera 1 at the 2nd, 3rd, 4th, 6th, 7th, and 8th frames can be filtered out from the total image set, and the second image set 2 can be added to the 3rd first image set.
[0063] Therefore, in this method, the second image set acquired by the target camera in the adjacent frames of the target frame can be screened out from the total image set, and the second image set can be added to the Nth first image set.
[0064] Method 2: In response to an acquisition parameter indicating that a candidate image in the N-1th candidate image set is acquired by a target camera at a target frame, a candidate camera that overlaps with a shooting range of the target camera is determined, a third image set acquired by the candidate camera at the target frame and an adjacent frame of the target frame is screened from the total image set, and the third image set is added to the Nth first image set.
[0065] It should be noted that the overlapping shooting ranges of the target camera and the candidate camera means that part or all of the shooting range of the target camera overlaps with the shooting range of the candidate camera. In some examples, the shooting range of the candidate camera includes the shooting range of the target camera, or the shooting range of the target camera includes the shooting range of the candidate camera.
[0066] It can be understood that the images captured by the candidate camera in the target frame and the adjacent frames of the target frame have a high similarity with the candidate images in the N-1th candidate image set.
[0067] It should be noted that there is no excessive limitation on the number of candidate cameras. For example, the number of candidate cameras can be 1, 3, etc.
[0068] For example, the second candidate image set includes candidate images A and B.
[0069] If the acquisition parameters of candidate image A indicate that candidate image A was acquired by target camera 1 at the 10th frame, candidate cameras 3 and 4 that coincide with the shooting range of target camera 1 can be determined. Then, a second image set 3 acquired by candidate camera 3 at the 7th, 8th, 9th, 10th, 11th, 12th, and 13th frames can be screened out from the total image set. A second image set 4 acquired by candidate camera 4 at the 7th, 8th, 9th, 10th, 11th, 12th, and 13th frames can also be screened out from the total image set, and the second image sets 3 and 4 are added to the third first image set.
[0070] If the acquisition parameters of candidate image B indicate that candidate image B was acquired by target camera 2 at the 5th frame, candidate camera 5 that overlaps with the shooting range of target camera 2 can be determined. Then, a second image set 5 acquired by candidate camera 5 at the 2nd, 3rd, 4th, 5th, 6th, 7th, and 8th frames can be screened out from the total image set, and the second image set 5 is added to the 3rd first image set.
[0071] In one embodiment, determining candidate cameras that overlap with the target camera's shooting range includes obtaining camera grouping results, where the shooting ranges of multiple cameras in the same group overlap, and determining the remaining cameras in the same group as the target camera as candidate cameras. Thus, the method can determine candidate cameras based on the camera grouping results.
[0072] It is understandable that multiple cameras with overlapping shooting ranges can be divided into the same group.
[0073] For example, the grouping results for cameras 1 to 5 include group 1 and group 2, where group 1 includes camera 1, camera 3, and camera 4, and group 2 includes camera 2 and camera 5. The shooting ranges of cameras 1, 3, and 4 overlap, and the shooting ranges of cameras 2 and 5 overlap.
[0074] Therefore, the method can determine the candidate camera that overlaps with the shooting range of the target camera, filter out the third image set captured by the candidate camera in the target frame and the adjacent frame of the target frame from the total image set, and add the third image set to the Nth first image set.
[0075] S204 : Retrieve an Nth candidate image set from the Nth first image set based on the target scene.
[0076] In one embodiment, retrieving an Nth candidate image set from an Nth first image set based on a target scene includes obtaining a target feature representation of a target object contained in the target scene, obtaining a first feature representation of each first image in the Nth first image set, and retrieving an Nth candidate image set from the Nth first image set based on the target feature representation and the first feature representation. Thus, the method can comprehensively consider the target feature representation and the first feature representation to retrieve an Nth candidate image set from the N first image sets.
[0077] It should be noted that there are no excessive restrictions on target objects, for example, target objects may include pedestrians, vehicles, traffic signs, trees, etc. There is no excessive restriction on the number of target objects, and a target scene may contain multiple target objects.
[0078] It is understandable that different target scenes may correspond to different target objects. For example, if the target scene includes a cargo scene, the target objects may include loaded objects, road signs, etc.; if the target scene includes a passenger scene, the target objects may include pedestrians, vehicles, etc.; if the target scene includes a cleaning scene, the target objects may include cleaned objects (such as furniture, walls, floors, etc.); if the target scene includes a monitoring scene, the target objects may include monitored objects (such as factory machinery).
[0079] In some examples, obtaining a target feature representation of a target object included in the target scene includes performing few-shot learning on a second sample image set of the target object based on a multimodal feature extraction model to obtain the target feature representation. The target feature representation may include a class center feature representation.
[0080] It is understandable that the number of second sample image sets of the target object may be small. In this method, a large model of multimodal feature extraction can be used to perform few-sample learning on the second sample image set of the target object to obtain target feature representation, thereby realizing feature extraction in few-sample scenarios.
[0081] In some examples, obtaining a first feature representation for each first image in the Nth first image set includes detecting a target region from the first image based on a general object detection large model, and extracting the first feature representation from the target region based on a multimodal feature extraction large model. It should be noted that the general object detection large model can be trained based on massive open source data, and the training process is not further limited herein.
[0082] In some examples, based on the target feature representation and the first feature representation, an Nth candidate image set is retrieved from the Nth first image set, including obtaining the similarity between the target feature representation and the first feature representation, and adding images with a similarity greater than a set threshold to the Nth candidate image set.
[0083] S205 : Obtain a first sample image set corresponding to the target scene based on the M candidate image sets.
[0084] S206: Training the autonomous driving model based on the first sample image set.
[0085] The relevant contents of steps S205-S206 can be found in the above embodiment and will not be repeated here.
[0086] In summary, according to the training method of the autonomous driving model of the embodiment of the present disclosure, based on the N-1th candidate image set, the Nth first image set is screened out from the total image set, and based on the target scene, the Nth candidate image set is retrieved from the Nth first image set. The N-1th candidate image set and the target scene can be comprehensively considered to retrieve the Nth candidate image set, thereby improving the accuracy of the Nth candidate image set.
[0087] Figure 3 3 is a flowchart of a method for training an autonomous driving model according to the third embodiment of the present disclosure.
[0088] like Figure 3 As shown, the training method of the autonomous driving model of the third embodiment of the present disclosure includes:
[0089] S301, obtaining a target scene for the autonomous driving model to be trained.
[0090] S302 , in the case of the first round of search, obtaining the grouping result of the cameras, wherein the shooting ranges of multiple cameras in the same group overlap.
[0091] The relevant contents of steps S301-S302 can be found in the above embodiment and will not be repeated here.
[0092] S303: extract a set number of cameras from each group.
[0093] S304 , extracting images from the images captured by the camera according to the set frame extraction frequency, and adding the extracted images to the first first image set.
[0094] It should be noted that the total image set includes images captured by each camera.
[0095] It should be noted that there are no excessive restrictions on the set number and the set frame extraction frequency. For example, the set number can be 1, 2, etc., and the set frame extraction frequency can be 1 frame every 3 seconds, 1 frame every 15 seconds, etc.
[0096] It is understandable that different cameras may correspond to the same set frame rate or different set frame rates.
[0097] For example, the grouping results of cameras 1 to 5 include group 1 and group 2, where group 1 includes camera 1, camera 3, and camera 4, and group 2 includes camera 2 and camera 5.
[0098] Camera 1 may be extracted from group 1, and an image may be extracted from the images captured by camera 1 at a set frame extraction frequency of 1 frame every 3 seconds, and the extracted image may be added to the first first image set.
[0099] Camera 2 may be extracted from group 2, and an image may be extracted from the images captured by camera 2 at a set frame extraction frequency of 1 frame every 15 seconds, and the extracted image may be added to the first first image set.
[0100] S305 : Retrieve a first candidate image set from the first first image set based on the target scene.
[0101] For the relevant content of step S305, please refer to the relevant content of step S204, which will not be repeated here.
[0102] S306 , in the case of the Nth round of retrieval, based on the N-1th candidate image set retrieved in the N-1th round and the target scene, retrieve the Nth candidate image set from the total image set, where 2≤N≤M, and M is an integer greater than 1.
[0103] S307 : Obtain a first sample image set corresponding to the target scene based on the M candidate image sets.
[0104] S308: Training the autonomous driving model based on the first sample image set.
[0105] The relevant contents of steps S306-S308 can be found in the above embodiment and will not be repeated here.
[0106] In summary, according to the training method of the autonomous driving model of the embodiment of the present disclosure, a set number of cameras can be extracted in each group based on the grouping results of the cameras, and images can be extracted from the images captured by the extracted cameras according to the set frame extraction frequency, and the extracted images are added to the first first image set to achieve the screening of the first first image set from the total image set, and based on the target scene, the first candidate image set is retrieved from the first first image set.
[0107] like Figure 4 As shown, taking M=2 as an example, in the first round of retrieval, the first first image set is screened out from the total image set, the target area is detected from the first image in the first first image set based on the general target detection model, and the first feature representation is extracted from the target area based on the multimodal feature extraction model.
[0108] Based on a large multimodal feature extraction model, a few-shot learning is performed on a second sample image set of a target object contained in a target scene to obtain a target feature representation.
[0109] Based on the first feature representation and the target feature representation, a first candidate image set is retrieved from the first first image set.
[0110] In the second round of retrieval, based on the first candidate image set retrieved in the first round, the second first image set is screened out from the total image set. Based on the general target detection model, the target area is detected from the first image in the second first image set. Based on the multimodal feature extraction model, the first feature representation is extracted from the target area.
[0111] Based on the first feature representation and the target feature representation, a second candidate image set is retrieved from the second first image set.
[0112] A first sample image set corresponding to the target scene is selected from the first and second candidate image sets, and the autonomous driving model is trained based on the first sample image set.
[0113] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0114] According to an embodiment of the present disclosure, the present disclosure also provides a training device for an autonomous driving model, which is used to implement the above-mentioned training method for the autonomous driving model.
[0115] Figure 5 4 is a block diagram of a training device for an autonomous driving model according to a first embodiment of the present disclosure.
[0116] like Figure 5 As shown, the training device 500 of the autonomous driving model of an embodiment of the present disclosure includes: a first acquisition module 501, a retrieval module 502, a second acquisition module 503 and a training module 504.
[0117] A first acquisition module 501 is used to acquire a target scene to be trained for the autonomous driving model;
[0118] A retrieval module 502 is configured to retrieve a first candidate image set from the total image set based on the target scene in the first round of retrieval;
[0119] The retrieval module 502 is further configured to retrieve an Nth candidate image set from the total image set in the Nth round of retrieval based on the N-1th candidate image set retrieved in the N-1th round of retrieval and the target scene, where 2≤N≤M, and M is an integer greater than 1;
[0120] A second acquisition module 503 is configured to obtain a first sample image set corresponding to the target scene based on the M candidate image sets;
[0121] The training module 504 is used to train the autonomous driving model based on the first sample image set.
[0122] In one embodiment of the present disclosure, the retrieval module 502 is further used to: screen out the Nth first image set from the total image set based on the N-1th candidate image set; and retrieve the Nth candidate image set from the Nth first image set based on the target scene.
[0123] In one embodiment of the present disclosure, the retrieval module 502 is further configured to: filter out the Nth first image set from the total image set based on acquisition parameters of the candidate images in the N-1th candidate image set.
[0124] In one embodiment of the present disclosure, the retrieval module 502 is further configured to: in response to the acquisition parameter indicating that the candidate image in the N-1th candidate image set is acquired by the target camera at the target frame, filter out, from the total image set, a second image set acquired by the target camera at a frame adjacent to the target frame; and add the second image set to the Nth first image set.
[0125] In one embodiment of the present disclosure, the retrieval module 502 is further configured to: in response to the acquisition parameter indicating that the candidate image in the N-1th candidate image set was acquired by the target camera at the target frame, determine a candidate camera that overlaps with the shooting range of the target camera; filter out a third image set acquired by the candidate camera at the target frame and an adjacent frame of the target frame from the total image set; and add the third image set to the Nth first image set.
[0126] In one embodiment of the present disclosure, the retrieval module 502 is further configured to: obtain camera grouping results, wherein the shooting ranges of multiple cameras in the same group overlap; and determine the remaining cameras in the same group as the target camera as the candidate cameras.
[0127] In one embodiment of the present disclosure, the retrieval module 502 is further used to: obtain a target feature representation of a target object contained in the target scene; obtain a first feature representation of each first image in the Nth first image set; and retrieve the Nth candidate image set from the Nth first image set based on the target feature representation and the first feature representation.
[0128] In one embodiment of the present disclosure, the retrieval module 502 is further configured to: perform few-sample learning on a second sample image set of the target object based on a multimodal feature extraction large model to obtain the target feature representation.
[0129] In one embodiment of the present disclosure, the retrieval module 502 is further used to: obtain camera grouping results, wherein the shooting ranges of multiple cameras in the same group overlap; extract a set number of cameras in each group; extract images from the images captured by the extracted cameras according to a set frame extraction frequency, and add the extracted images to the first first image set; and retrieve the first candidate image set from the first first image set based on the target scene.
[0130] In summary, the training device for an autonomous driving model according to an embodiment of the present disclosure obtains a target scene for training the autonomous driving model. In the first round of retrieval, the first candidate image set is retrieved from the total image set based on the target scene. In the Nth round of retrieval, the Nth candidate image set is retrieved from the total image set based on the N-1th candidate image set retrieved in the N-1th round and the target scene, where 2≤N≤M, and M is an integer greater than 1. Based on the M candidate image sets, a first sample image set corresponding to the target scene is obtained, and the autonomous driving model is trained based on the first sample image set. Thus, the Nth round of retrieval relies on the N-1th candidate image set retrieved in the N-1th round, thereby achieving multiple rounds of iterative retrieval of the total image set. Each round of retrieval can obtain one candidate image set to obtain the first sample image set. Compared to the related art, which often only performs a single retrieval of the total image set, this reduces the difficulty of searching the total image set and improves the accuracy of image retrieval, thereby reducing the difficulty of training the autonomous driving model and improving the training accuracy of the autonomous driving model.
[0131] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0132] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0133] like Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0134] Multiple components in the electronic device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0135] The computing unit 601 may be a variety of general and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as Figures 1 to 4 The training method of the autonomous driving model described above. For example, in some embodiments, the training method of the autonomous driving model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the training method of the autonomous driving model described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the training method of the autonomous driving model in any other appropriate manner (for example, by means of firmware).
[0136] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0140] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0141] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0142] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the steps of the training method of the autonomous driving model described in the above embodiment of the present disclosure are implemented.
[0143] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0144] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training an autonomous driving model, comprising: Obtain the target scenario for the autonomous driving model to be trained; In the case of the first round of retrieval, based on the target scene, the first candidate image set is retrieved from the total image set; In the case of the Nth round of retrieval, based on the acquisition parameters of the candidate images of the N-1th candidate image set retrieved in the N-1th round, the Nth first image set is screened out from the total image set; Based on the target scene, retrieve an Nth candidate image set from the Nth first image set, where 2≤N≤M, and M is an integer greater than 1; Based on the M candidate image sets, obtaining a first sample image set corresponding to the target scene; Training the autonomous driving model based on the first sample image set; The retrieving the Nth candidate image set from the Nth first image set based on the target scene includes: Obtaining a target feature representation of a target object contained in the target scene; Obtaining a first feature representation of each first image in the Nth first image set; The Nth candidate image set is retrieved from the Nth first image set based on the target feature representation and the first feature representation.
2. The method according to claim 1, wherein The step of selecting the Nth first image set from the total image set based on the acquisition parameters of the candidate images in the N-1th candidate image set retrieved in the N-1th round includes: In response to the acquisition parameter indicating that the candidate image in the N-1th candidate image set is acquired by the target camera at a target frame, filtering out, from the total image set, a second image set acquired by the target camera at a frame adjacent to the target frame; The second image set is added to the Nth first image set.
3. The method according to claim 1, wherein The step of selecting the Nth first image set from the total image set based on the acquisition parameters of the candidate images in the N-1th candidate image set retrieved in the N-1th round includes: In response to the acquisition parameter indicating that the candidate image in the N-1th candidate image set is acquired by a target camera at a target frame, determining a candidate camera that overlaps with a shooting range of the target camera; Filtering out a third image set acquired by the candidate camera during the target frame and an adjacent frame of the target frame from the total image set; The third image set is added to the Nth first image set.
4. The method according to claim 3, characterized in that The determining of candidate cameras that overlap with the shooting range of the target camera includes: Obtain camera grouping results, where the shooting ranges of multiple cameras in the same group overlap; The remaining cameras in the same group as the target camera are determined as the candidate cameras.
5. The method according to claim 1, wherein The obtaining of a target feature representation of a target object contained in the target scene includes: Based on a multimodal feature extraction large model, a few-sample learning is performed on a second sample image set of the target object to obtain the target feature representation.
6. The method according to any one of claims 1 to 5, characterized in that The step of retrieving a first candidate image set from the total image set based on the target scene includes: Obtain camera grouping results, where the shooting ranges of multiple cameras in the same group overlap; A set number of cameras are sampled within each group; Extracting images from the images captured by the camera according to the set frame extraction frequency, and adding the extracted images to the first first image set; Based on the target scene, the first candidate image set is retrieved from the first first image set.
7. A training device for an autonomous driving model, comprising: The first acquisition module is used to obtain the target scene to be trained for the autonomous driving model; A retrieval module, configured to retrieve a first candidate image set from the total image set based on the target scene in the first round of retrieval; The retrieval module is further configured to, in an Nth round of retrieval, screen an Nth first image set from the total image set based on acquisition parameters of candidate images of the N-1th candidate image set retrieved in the N-1th round, and retrieve an Nth candidate image set from the Nth first image set based on the target scene, where 2≤N≤M, and M is an integer greater than 1; A second acquisition module is configured to obtain a first sample image set corresponding to the target scene based on the M candidate image sets; a training module, configured to train the autonomous driving model based on the first sample image set; Wherein, the retrieval module is further used to: Obtaining a target feature representation of a target object contained in the target scene; Obtaining a first feature representation of each first image in the Nth first image set; The Nth candidate image set is retrieved from the Nth first image set based on the target feature representation and the first feature representation.
8. The device according to claim 7, wherein The retrieval module is further used to: In response to the acquisition parameter indicating that the candidate image in the N-1th candidate image set is acquired by the target camera at a target frame, filtering out, from the total image set, a second image set acquired by the target camera at a frame adjacent to the target frame; The second image set is added to the Nth first image set.
9. The device according to claim 7, wherein The retrieval module is further used to: In response to the acquisition parameter indicating that the candidate image in the N-1th candidate image set is acquired by a target camera at a target frame, determining a candidate camera that overlaps with a shooting range of the target camera; Filtering out a third image set acquired by the candidate camera during the target frame and an adjacent frame of the target frame from the total image set; The third image set is added to the Nth first image set.
10. The device according to claim 9, characterized in that The retrieval module is further used to: Obtain camera grouping results, where the shooting ranges of multiple cameras in the same group overlap; The remaining cameras in the same group as the target camera are determined as the candidate cameras.
11. The device according to claim 7, wherein The retrieval module is further used to: Based on a multimodal feature extraction large model, a few-sample learning is performed on a second sample image set of the target object to obtain the target feature representation.
12. The device according to any one of claims 7 to 11, characterized in that The retrieval module is further used to: Obtain camera grouping results, where the shooting ranges of multiple cameras in the same group overlap; A set number of cameras are sampled within each group; Extracting images from the images captured by the camera according to the set frame extraction frequency, and adding the extracted images to the first first image set; Based on the target scene, the first candidate image set is retrieved from the first first image set.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the autonomous driving model as described in any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the training method of the autonomous driving model as described in any one of claims 1 to 6.
15. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the training method for an autonomous driving model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method, device and equipment and storage medium
CN112712138A
Image retrieval method and apparatus, storage medium, and device
WO2021159769A1