Generation method and device of perception model training data, equipment, vehicle and medium
By calculating the comprehensive score of the video generation model and selecting the appropriate model to generate perception model training data, the problems of high acquisition cost and insufficient scene coverage are solved, high-realism driving scene video generation is achieved, and the perception ability of the intelligent driving system is improved.
Patent Information
- Application Number
- CN202510782147.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-16
AI Technical Summary
In existing technologies, the acquisition cost of perception model training data is high, the scene coverage is insufficient, and the authenticity is low, making it difficult to effectively train the perception model of the intelligent driving system.
By obtaining evaluation data of user-generated tasks and video generation models, calculating the comprehensive score, and selecting the most appropriate video generation model, we can generate diverse and highly realistic driving scene videos as training data for the perception model.
It improves the coverage and richness of driving scene videos, reduces the cost of real data collection, and enhances the perception capabilities of intelligent driving systems.
Smart Images

Figure CN120656016A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data generation technology, and in particular to a method, device, equipment, vehicle and medium for generating perception model training data. Background Art
[0002] Intelligent driving systems rely on vast amounts of real-world data to train their perception models, enabling them to accurately identify obstacles and other traffic participants. These perception models are a core component of intelligent driving systems, determining whether a vehicle can safely and efficiently navigate complex driving environments.
[0003] In related technologies, driving scene videos used as perception training data mainly come from data collected from actual vehicles and a small amount of driving scene videos generated using neural network technology. However, these methods have the following major problems:
[0004] 1. High cost: Collecting high-quality real-world driving scene videos requires a lot of time and resources, especially when collecting data under different weather conditions, road types, and traffic conditions; Insufficient scene coverage: 2. Certain extreme or edge cases (such as events occurring under conditions such as heavy rain, heavy snow, and low light at night) occur less frequently in real life, making it difficult to obtain sufficient samples through field collection; 3. Low realism: Driving scene videos generated using neural networks often lack sufficient details and diversity, and cannot fully simulate the complex and changeable real-world driving scenarios. Summary of the Invention
[0005] The present application provides a method, apparatus, device, vehicle, and medium for generating perception model training data to solve the problems in related technologies such as high acquisition cost, insufficient scene coverage, and low authenticity of driving scene videos used as perception model training data.
[0006] The first aspect of the present application provides a method for generating perceptual model training data, comprising the following steps: obtaining user evaluation data on a current generation task of the perceptual model training data and historical generation tasks performed by multiple video generation models; extracting required features and multiple evaluation indicators in the current generation task, wherein the required features include at least one of video theme, video duration, style preference, picture quality requirements and specific elements, and the evaluation indicators include at least one of video quality, content matching, generation efficiency and user feedback score; calculating comprehensive scores of multiple video generation models based on multiple evaluation indicators and evaluation data, determining a target video generation model based on the comprehensive scores, and the target video generation model generating training data for the perceptual model based on the required features.
[0007] According to the above-mentioned technical means, the embodiment of the present application can extract the element features and multiple evaluation indicators of the user's current generation task, and calculate the comprehensive scores of multiple video generation models based on the evaluation data of historical generation tasks performed by multiple evaluation indicators and multiple video generation models, and then match the most suitable target video generation model according to the comprehensive score, and use the target video generation model to generate training data of the perception model that matches the current generation task based on the required features, thereby realizing the efficient generation of diversified and highly realistic driving scene videos through the video generation model, and using the driving scene videos as training data for the perception model, thereby improving the coverage and richness of the driving scenes in the vehicle perception model training data set, and reducing the cost of using real collected data to train the vehicle-side perception model.
[0008] Optionally, calculating the comprehensive scores of multiple video generation models based on multiple evaluation indicators and evaluation data also includes: obtaining the historical priorities of multiple evaluation indicators in historical generation tasks; obtaining the current priorities of multiple evaluation indicators in the current generation task; determining the weights of multiple evaluation indicators in the current generation task based on at least one of the current priority and the historical priority; extracting the scores of multiple evaluation indicators in the evaluation data; and calculating the comprehensive score of each video generation model based on the weights and scores of multiple evaluation indicators.
[0009] According to the above-mentioned technical means, the embodiment of the present application can determine the weights of multiple evaluation indicators in the current video generation task based on the current priorities of multiple evaluation indicators in the current video generation task process and at least one of the historical priorities of multiple evaluation indicators in the historical video generation task, and extract the scores of multiple evaluation indicators in the evaluation data, calculate the comprehensive score of each video generation model based on the scores and weights, and select the most matching target video generation model from multiple video generation models based on the comprehensive score to improve the authenticity, accuracy and quality of driving scene video generation.
[0010] Optionally, the comprehensive score is calculated as:
[0011] S i =ω1×Q i +ω2×C i +ω3×E i +ω4×U i ;
[0012] Among them, S i is the comprehensive score of the video generation model, ω1, ω2, ω3, and ω4 are the weights of video quality, content matching, generation efficiency, and user feedback score, respectively; Q i 、C i 、E i 、U i Corresponding to the video generation model M iThe scores on each evaluation indicator.
[0013] Optionally, extracting scores of multiple evaluation indicators from the evaluation data includes: performing quantization processing on the evaluation data to obtain scores of the multiple evaluation indicators.
[0014] According to the above technical means, the embodiment of the present application can quantify multiple evaluation data to obtain scores of multiple evaluation indicators, so as to provide a reliable selection basis in the subsequent video generation model selection process.
[0015] Optionally, before determining the target video generation model based on the comprehensive score, the method further includes: if the comprehensive score of each video generation model is lower than a preset score, reminding the user to adjust the current generation task or update the version of each video generation model.
[0016] According to the above technical means, the embodiment of the present application can remind users to adjust their requirements or update the video generation model version when the comprehensive score of all video generation models is lower than the preset score, thereby ensuring the quality of the driving scene video generation results and avoiding the failure to meet the user's video generation requirements.
[0017] Optionally, it also includes: obtaining current evaluation data of the current generation task; and using the current evaluation data to update the target video generation model to execute the evaluation data of the historical video generation task.
[0018] According to the above-mentioned technical means, the embodiment of the present application can continuously optimize the video generation model by collecting the evaluation data of the target video generation model in the current generation task process and updating the evaluation data of the historical generation tasks of the target video generation model, so as to better cope with future video generation tasks.
[0019] The second aspect of the present application provides a device for generating perception model training data, including: an acquisition module for obtaining user evaluation data on a current generation task of the perception model training data and historical generation tasks performed by multiple video generation models; an extraction module for extracting required features and multiple evaluation indicators in the current generation task, wherein the required features include at least one of video theme, video duration, style preference, picture quality requirements and specific elements, and the evaluation indicators include at least one of video quality, content matching, generation efficiency and user feedback score; a generation module for calculating the comprehensive scores of multiple video generation models based on multiple evaluation indicators and evaluation data, determining a target video generation model based on the comprehensive score, and the target video generation model generating training data for the perception model based on the required features.
[0020] Optionally, the generation module is further used to: obtain the historical priorities of multiple evaluation indicators in historical generation tasks; obtain the current priorities of multiple evaluation indicators in the current generation task; determine the weights of multiple evaluation indicators in the current generation task based on at least one of the current priority and the historical priority; extract the scores of multiple evaluation indicators in the evaluation data; and calculate the comprehensive score of each video generation model based on the weights and scores of multiple evaluation indicators.
[0021] Optionally, the comprehensive score is calculated as:
[0022] S i =ω1×Q i +ω2×C i +ω3×E i +ω4×U i ;
[0023] Among them, S i is the comprehensive score of the video generation model, ω1, ω2, ω3, and ω4 are the weights of video quality, content matching, generation efficiency, and user feedback score, respectively; Q i 、C i 、E i 、U i Corresponding to the video generation model M i The scores on each evaluation indicator.
[0024] Optionally, the generation module is further used to: perform quantitative processing on the evaluation data to obtain scores of multiple evaluation indicators.
[0025] Optionally, it also includes: a reminder module, which is used to remind the user to adjust the current generation task or update the version of each video generation model if the comprehensive score of each video generation model is lower than a preset score before determining the target video generation model based on the comprehensive score.
[0026] Optionally, it also includes: an updating module for obtaining current evaluation data of the current generation task; and using the current evaluation data to update the evaluation data of the target video generation model to perform historical video generation tasks.
[0027] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein the processor executes the program to perform a method for generating perceptual model training data as described in the above embodiment.
[0028] A fourth aspect of the present application provides a vehicle having a perception model provided thereon, wherein the perception model is trained using training data generated by the method for generating perception model training data as described in the above-mentioned embodiment.
[0029] The fifth aspect of the present application provides a computer-readable storage medium on which a computer program or instruction is stored. The computer program or instruction is executed by a processor to perform the method for generating edge-aware model training data as described in the above embodiment.
[0030] The sixth aspect of the present application provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed, they can implement the method for generating perception model training data as described in the above embodiment.
[0031] Therefore, this application has at least the following beneficial effects:
[0032] The embodiment of the present application can extract the element features and multiple evaluation indicators of the user's current generation task, and calculate the comprehensive scores of multiple video generation models based on the evaluation data of the multiple evaluation indicators and the historical generation tasks performed by multiple video generation models. Then, the most appropriate target video generation model is matched according to the comprehensive score, and the target video generation model is used to generate training data for the perception model that matches the current generation task based on the required features, thereby achieving efficient generation of diverse and highly realistic driving scene videos through the video generation model, and using the driving scene videos as training data for the perception model, thereby improving the coverage and richness of driving scenes in the vehicle perception model training data set, and reducing the cost of using real collected data to train the vehicle-side perception model. As a result, technical problems in the related art such as the high acquisition cost of driving scene videos used as perception model training data, insufficient scene coverage, and low authenticity are solved.
[0033] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0035] Figure 1 A flowchart of a method for generating perception model training data according to an embodiment of the present application;
[0036] Figure 2 Flowchart of the driving scene video generation process and training perception model provided in accordance with an embodiment of the present application;
[0037] Figure 3 A schematic diagram of a driving scene video generation process provided according to a specific embodiment of the present application;
[0038] Figure 4 A flowchart of a driving scene video generated based on real vehicle-side data according to an embodiment of the present application;
[0039] Figure 5 This is an example diagram of a device for generating perception model training data according to an embodiment of the present application;
[0040] Figure 6 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0041] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0042] In related technologies, perception training data mainly comes from data actually collected by vehicles and a small amount of data generated by using neural network (such as NeRF) + material library permutation and combination + game engine technology. However, the cost of actually collecting data is high and the scene coverage is insufficient. The shortcomings of the method of generating data using neural network + material library permutation and combination + game engine technology are: NeRF is a non-generative AI, which is distorted and violates the objective laws of reality. It also requires manual addition of materials, such as adding generalized scenes such as rainy and foggy environments, and there is a problem that the added materials are not well matched with the real scenes.
[0043] To this end, the present application provides a method for generating perception model training data, in which an evaluation strategy for multiple video generation models can be constructed, wherein the evaluation strategy is based on the evaluation data of the current video generation task and the historical video generation tasks performed by multiple video generation models to select the most matching target video generation model from the multiple video generation models, and construct an intelligent agent based on the evaluation data and the evaluation strategy, thereby realizing the use of the intelligent agent to select the target video generation model for the current video generation task, and generating training data for the perception model based on the target video generation model. The video generation model can be used to efficiently generate diverse and highly realistic driving scene videos, and the driving scene videos can be used as training data for the perception model, thereby improving the coverage and richness of driving scenes in the vehicle perception model training data set, and reducing the cost of using real collected data to train the vehicle-side perception model. As a result, the problems of high acquisition cost, insufficient scene coverage and low authenticity of driving scene videos used as perception model training data in the related art are solved.
[0044] Specifically, Figure 1 A flowchart of a method for generating perception model training data provided in an embodiment of the present application.
[0045] It should be noted that the method for generating perception model training data in this application mainly uses the video generation model to generate edge (Conner Case) scene videos, where edge scenes refer to those specific driving scenarios that have a low probability of occurrence and are less common, but are extremely critical to the perception capabilities of the intelligent driving system, such as driving scenarios in extreme weather conditions, complex traffic conditions or special circumstances.
[0046] like Figure 1 As shown, the method for generating the perception model training data includes the following steps:
[0047] In step S101, user evaluation data on a current generation task of perception model training data and historical generation tasks performed by multiple video generation models are obtained.
[0048] Among them, the video generation models can be Sora, Pika, Stable Diffusion, etc.; the evaluation data is the performance data of multiple video generation models performing video generation tasks. The evaluation data can include multiple dimensions, such as picture quality, color accuracy, etc., which are mainly summarized into dimensions such as video quality, content matching, generation efficiency, and user feedback scores. Video quality and generation efficiency can be obtained based on automated evaluation methods, content matching can be obtained based on manual evaluation or automated evaluation based on text semantic analysis, and user feedback scores are based on manual evaluation.
[0049] Specifically, video quality: average resolution, color accuracy, picture smoothness, and the presence of picture defects (such as noise, blur, etc.);
[0050] Content matching: The degree of fit between the generated video content and the input theme and element requirements can be calculated through manual evaluation or automated evaluation based on text semantic analysis to determine the matching accuracy.
[0051] Generation efficiency: The average time required to generate a video of a certain length, including the entire time from receiving the instruction to outputting a complete and usable video;
[0052] User feedback rating: Collect the comprehensive ratings given by users who use the video generation model to generate videos (you can set ratings in multiple dimensions, such as authenticity, practicality, etc., and then summarize them into a comprehensive rating).
[0053] It is understandable that the embodiments of the present application can obtain evaluation data for the current generation task of the perception model training data and the historical generation tasks performed by multiple video generation models.
[0054] In step S102, required features and multiple evaluation indicators in the current generation task are extracted, wherein the required features include at least one of video theme, video length, style preference, image quality requirements and specific elements, and the evaluation indicators include at least one of video quality, content matching, generation efficiency and user feedback score.
[0055] Among them, the image quality requirements in the element features may include resolution, color style, etc., and specific elements may include specific characters, driving scenes, etc.
[0056] It is understandable that the embodiments of the present application can extract the required features and multiple evaluation indicators in the current generation task in order to subsequently determine the target video generation model.
[0057] In step S103, comprehensive scores of multiple video generation models are calculated based on multiple evaluation indicators and evaluation data, and a target video generation model is determined based on the comprehensive scores. The target video generation model generates training data of the perception model based on the required features.
[0058] Among them, the target video generation model is the video generation model that best matches the current generation task.
[0059] It can be understood that the embodiments of the present application can calculate the comprehensive scores of multiple video generation models based on multiple evaluation indicators and evaluation data, and determine the target video generation model that best matches the current generation task based on the comprehensive score. The target video generation model generates training data for the perception model based on the required features, and generates diversified driving scene videos through the video generation model. Compared with driving scene videos constructed by real-life scene data, the cost is lower and the coverage is wider. Compared with the controllability and authenticity of scene videos generated by neural networks, the generated driving scene videos are used as training data for the perception model to improve the vehicle's perception capabilities.
[0060] In an embodiment of the present application, the comprehensive scores of multiple video generation models are calculated based on multiple evaluation indicators and evaluation data, and also include: obtaining the historical priorities of multiple evaluation indicators in historical generation tasks; obtaining the current priorities of multiple evaluation indicators in the current generation task; determining the weights of multiple evaluation indicators in the current generation task based on at least one of the current priority and the historical priority; extracting the scores of multiple evaluation indicators in the evaluation data; and calculating the comprehensive score of each video generation model based on the weights and scores of multiple evaluation indicators.
[0061] It can be understood that the embodiment of the present application can determine the weights of multiple evaluation indicators in the current video generation task based on the current priorities of multiple evaluation indicators in the current generation task and at least one of the historical priorities of multiple evaluation indicators in the historical video generation task, and extract the scores of multiple scoring indicators in the evaluation data, and then calculate the comprehensive score of each video generation model based on the weights and scores of multiple evaluation indicators.
[0062] Among them, the specific steps for determining the weights of multiple evaluation indicators are: judging whether the user has specified the priority of the evaluation indicators in the current generation task, that is, the focus of the user's greatest concern in generating the video; if multiple evaluation indicators have specified the priority of each evaluation indicator when the current generation task is generated, then the weights of the multiple evaluation indicators in the current generation task are determined according to the priority; otherwise, as long as there is an evaluation indicator without priority, for example, the user only specifies that the video quality is the most important concern, but does not specify other requirements, then the weights of the multiple evaluation indicators in the current generation task can be determined based on the priority of the multiple evaluation indicators in the historical generation task and the priority of the evaluation indicators specified in the current generation task.
[0063] For example, if before executing the current generation task, the user specifies that the first focus is video quality, the second focus is content matching, the third focus is generation efficiency, and the fourth focus is user feedback score, then the weight of video quality can be set to 0.6, the weight of content matching can be set to 0.3, the weight of generation efficiency can be set to 0.08, and the weight of user feedback score can be set to 0.02. If the user only specifies to focus on video quality, the weight of video quality can be set to 0.7. The weights of other evaluation indicators can be determined according to the order of evaluation indicators that historical generation tasks mainly focus on. For example, based on the analysis of historical generation tasks, it is found that users generally focus on video quality first, then generation efficiency, and then content matching, and finally user feedback score. The weight of generation efficiency can be set to 0.2, the weight of content matching can be set to 0.06, and the weight of user feedback score can be set to 0.04.
[0064] In addition, since the user feedback score in the evaluation index is a subjective score of the user, the user feedback score can be set as a fixed weight in advance to improve the selection accuracy of the target video generation model.
[0065] In an embodiment of the present application, extracting scores of multiple evaluation indicators from the evaluation data includes: performing quantization processing on the evaluation data to obtain scores of the multiple evaluation indicators.
[0066] It is understandable that the embodiment of the present application can quantify the evaluation data of each video generation model to obtain scores of multiple evaluation indicators.
[0067] For example, the video quality index is set with a specific score range according to different indicators (such as 8-10 points for a resolution of 1080p or above, 5-7 points for 720p-1080p, etc.), the content matching degree is expressed as a percentage, the generation efficiency is expressed in quantitative forms such as the average length of video generated per minute, and the user feedback score directly adopts the existing score value, and the total weight of multiple evaluation indicators is 1.
[0068] For example, the video generation model M1 scores 5 points in content matching, while the video generation model M2 scores 8 points in content matching.
[0069] In the embodiment of the present application, the calculation formula of the comprehensive score is:
[0070] S i =ω1×Q i +ω2×C i +ω3×E i +ω4×U i ;
[0071] Among them, S i is the comprehensive score of the video generation model, ω1, ω2, ω3, and ω4 are the weights of video quality, content matching, generation efficiency, and user feedback score, respectively; Q i 、C i 、E i 、U i Corresponding to the video generation model M i The scores on each evaluation indicator.
[0072] In an embodiment of the present application, before determining the target video generation model based on the comprehensive score, it also includes: if the comprehensive score of each video generation model is lower than the preset score, the user is reminded to adjust the current generation task or update the version of each video generation model.
[0073] Among them, the preset score can be set according to the specific situation, and there is no specific limitation for comparison, such as 90 points.
[0074] It is understandable that, in an embodiment of the present application, when the total evaluation score of each video generation model is lower than a preset score, the user is prompted to adjust the current generation task or update the version of each video generation model to improve the quality of the driving scene video generation results.
[0075] In an embodiment of the present application, it also includes: obtaining current evaluation data of the current generation task; and using the current evaluation data to update the target video generation model to execute the evaluation data of the historical video generation task.
[0076] It can be understood that the embodiment of the present application can obtain the evaluation data of the target video generation model performing the current generation task after the target video generation model completes the current generation task, that is, after generating the edge driving scene video that meets the task requirements of the current generation task, and use the current evaluation data to update the evaluation data of the historical generation tasks of the target video generation model to optimize the target video generation model and better adapt to the requirements of selecting the optimal model for video generation in different scenarios.
[0077] Specifically, for example, according to the evaluation data of the target video generation model in the historical generation task, the target video generation model's content matching score is 9 points, but after this generation task, the target video generation model's content matching score is 6 points. Then, the target video generation model's content matching score can be lowered based on the 6 points this time, for example, it can be lowered to 8.5 points.
[0078] The following describes a specific embodiment of the intelligent agent evaluating the video generation model in this application to select the most appropriate video generation model to generate a driving scene video.
[0079] For example, a user needs a video containing a rainy city driving scene. The specific video requirements entered by the user include weather: rainy day, location: city, duration: 5 minutes, and image quality: HD (1080p), and the user specifies priority attention to content matching. The weight of each video generation model on different evaluation indicators is set according to the most important evaluation indicators that the user pays attention to in this video generation task, and the priority of the evaluation indicators that users generally pay attention to in historical video generation tasks. In this video generation task, the weight of content matching is set to the highest, and the weights of other evaluation indicators are set according to the priority of the evaluation indicators that users generally pay attention to in historical video generation tasks. The video generation model with the highest total evaluation score is calculated as the target video generation model. The target video generation model generates a video containing a rainy city driving scene based on the specific video requirements including weather: rainy day, location: city, duration: 5 minutes, and image quality: HD (1080p).
[0080] In addition, it should be noted that the embodiment of the present application can also determine the generalized prompts for driving scene video generation based on the real perception video transmitted back by the vehicle and the content in the real perception video. The video generation model generates more driving scene videos in combination with the real perception video transmitted back by the vehicle, thereby covering more possible scenarios, including some less common but critical edge cases, and the perception model can be trained based on the real perception video transmitted back by the vehicle and the edge driving scene video generated by this application to improve the perception ability of the perception model.
[0081] Specifically, the method for generating the perception model training data and the process of using the training data to train the vehicle-side perception model in the embodiment of the present application are as follows: Figure 2 As shown, the video generation model is hereinafter referred to as the large model, which includes the following steps:
[0082] 1. Pre-training the agent.
[0083] 1.1 Data collection.
[0084] 1.1.1 Collect video generation requirements.
[0085] (1) Collecting users' specific requirements for the videos to be generated (i.e., as training data for the perception training model), including but not limited to the video theme, video length, style preferences, image quality requirements (resolution, color style, etc.), and whether there are specific elements (such as specific characters, scenes, etc.);
[0086] (2) Organize the above requirements into prompts that are easy for large models to understand.
[0087] 1.1.2 Collect data related to the large model.
[0088] (2) Input the above video requirements to each participating large model to generate a video;
[0089] (2) Regularly collect performance data (i.e., evaluation data) from past video generation tasks, for example:
[0090] Generated video quality: average resolution, color accuracy, picture smoothness, and the presence of picture defects (such as noise, blur, etc.);
[0091] Content matching: The degree of fit between the generated video content and the input theme and element requirements can be calculated through manual evaluation or automated evaluation based on text semantic analysis to determine the matching accuracy.
[0092] Generation efficiency: The average time required to generate a video of a certain length, including the entire time from receiving the instruction to outputting a complete and usable video;
[0093] User feedback rating: Collect the comprehensive ratings given by users who use the large model to generate videos (you can set ratings in multiple dimensions, such as authenticity, practicality, etc., and then summarize them into a comprehensive rating).
[0094] 1.2 Feature extraction and quantization.
[0095] 1.2.1 Video generation requires feature extraction.
[0096] The video requirement description entered by the user is parsed, key features are extracted, and quantified. For example, video themes are categorized and assigned category numbers, style preferences are assigned corresponding numerical codes, and numerical requirements such as duration are directly extracted, allowing for rapid matching with large models for subsequent video generation.
[0097] 1.2.2 Large model generation feature extraction.
[0098] Based on the collected data on the past performance of the large model, corresponding key features are extracted and quantified. For example, the generated video quality indicators are set to specific score ranges according to different indicators (such as 8-10 points for resolutions reaching 1080p and above, 5-7 points for 720p-1080p, etc.), the content matching degree is expressed as a percentage, and the generation efficiency is expressed in quantitative forms such as the average length of the generated video per minute. User feedback scores are directly based on existing rating values.
[0099] 1.3 Construct evaluation indicators.
[0100] 1.3.1 Determine the weight.
[0101] We assign appropriate weights to different evaluation metrics based on the focus of the video generation task and the priorities of users. For example, if video quality is the most critical, we can assign it a higher weight (e.g., 0.4), content matching a weight of 0.3, generation efficiency a weight of 0.2, and user feedback a weight of 0.1 (the sum of the weights should be 1).
[0102] 1.3.2 Construct a comprehensive evaluation formula.
[0103] Based on the above weights and the extracted quantified features, a mathematical formula for comprehensive evaluation of the large model is constructed. For example, let the large model be M i , comprehensive evaluation score S i It can be expressed as:
[0104] S i =ω1×Q i +ω2×C i +ω3×E i +ω4×U i ;
[0105] Among them, S i is the comprehensive score of the large model, ω1, ω2, ω3, and ω4 are the weights of video quality, content matching, generation efficiency, and user feedback score respectively; Q i 、C i 、E i 、U i Corresponding to the large model M i Quantified scores on each indicator.
[0106] 1.4 Model selection
[0107] 1.4.1 Calculate the comprehensive score of each large model.
[0108] Substitute the quantitative characteristic values corresponding to each large model involved in the selection into the comprehensive evaluation formula, and calculate their comprehensive evaluation scores S1, S2, S3... respectively.
[0109] 1.4.2 Select the optimal model.
[0110] Compare the comprehensive evaluation scores of each large model and select the one with the highest score as the optimal choice for the video generation task. At the same time, a threshold can be set. If all models score below this threshold (preset score), the user may need to adjust their requirements or consider other additional factors (such as whether to update the large model version).
[0111] 2. The specific process of using the intelligent agent to generate driving scene videos is as follows: Figure 3 shown.
[0112] 2.1 Input the video generation requirements to the intelligent agent obtained by the above training, and select the appropriate large model to generate the video.
[0113] 2.2 Store the generated videos as a dataset for perceptual model training.
[0114] 2.3 Feedback and Updates.
[0115] 2.3.1 Collect feedback.
[0116] After selecting the large model to generate the video, we collect feedback on the actual situation during the generation of the driving scene video, such as the deviation between the actual generated video and the requirements, whether new problems arise (such as freezing, unreasonable content, etc.), and the user's final satisfaction score for the generated driving scene video.
[0117] 2.3.2 Update model data.
[0118] Based on the newly collected feedback, the performance data of the corresponding large model is updated to facilitate more accurate model selection decisions. This allows the agent to continuously learn and optimize, adapting to the optimal model selection requirements for video generation in different scenarios.
[0119] 3. Generalization, that is, generating driving scene videos based on real data from the vehicle side. The process is as follows Figure 4 shown.
[0120] 3.1 The data sent back by the vehicle is used to generate the required video by using the generative model's ability to support image or video generation and combining it with generalized prompts;
[0121] 3.2 Use the real data sent back by the vehicle and the generated data to train the perception model and continuously iterate the model capabilities.
[0122] In summary, the embodiments of the present application can support multiple video generation large models through the intelligent agent, automatically select the appropriate large model to generate diversified and highly realistic driving scene videos, and the driving scene video generation cost is lower. It can also collect video feedback generated each time to continuously optimize the intelligent agent, generate better driving scene videos, and use the driving scene video as a training data set to train the vehicle perception model, thereby improving the coverage of the scene in the perception model training, solving the problem of intelligent driving vehicles recognizing general obstacles, and improving the capabilities of the intelligent driving perception terminal.
[0123] According to the method for generating perception model training data proposed in the embodiment of the present application, the element features and multiple evaluation indicators of the user's current generation task can be extracted, and the comprehensive scores of multiple video generation models can be calculated based on the multiple evaluation indicators and the evaluation data of historical generation tasks performed by multiple video generation models. Then, the most suitable target video generation model is matched according to the comprehensive score, and the target video generation model is used to generate training data of the perception model that matches the current generation task based on the required features, thereby realizing efficient generation of diversified and highly realistic driving scene videos through the video generation model, and using the driving scene videos as training data for the perception model, thereby improving the coverage and richness of driving scenes in the vehicle perception model training data set, and reducing the cost of using real collected data to train the vehicle-side perception model.
[0124] Next, a device for generating perception model training data according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0125] Figure 5 It is a block diagram of a device for generating perceptual model training data according to an embodiment of the present application.
[0126] like Figure 5 As shown, the device 10 for generating perceptual model training data includes: an acquisition module 100, an extraction module 200 and a generation module 300.
[0127] Among them, the acquisition module 100 is used to obtain the user's evaluation data on the current generation task of the perception model training data and the historical generation tasks performed by multiple video generation models; the extraction module 200 is used to extract the required features and multiple evaluation indicators in the current generation task, wherein the required features include video theme, video length, style preference, picture quality requirements and at least one of specific elements, and the evaluation indicators include video quality, content matching, generation efficiency and user feedback score; the generation module 300 is used to calculate the comprehensive scores of multiple video generation models based on multiple evaluation indicators and evaluation data, determine the target video generation model based on the comprehensive score, and the target video generation model generates training data for the perception model based on the required features.
[0128] In the implementation of this application, the generation module 300 is further used to: obtain the historical priorities of multiple evaluation indicators in historical generation tasks; obtain the current priorities of multiple evaluation indicators in the current generation task; determine the weights of multiple evaluation indicators in the current generation task based on at least one of the current priority and the historical priority; extract the scores of multiple evaluation indicators in the evaluation data; and calculate the comprehensive score of each video generation model based on the weights and scores of multiple evaluation indicators.
[0129] In the implementation of this application, the calculation formula for the comprehensive score is:
[0130] S i =ω1×Q i +ω2×C i +ω3×E i +ω4×U i ;
[0131] Among them, S i is the comprehensive score of the video generation model, ω1, ω2, ω3, and ω4 are the weights of video quality, content matching, generation efficiency, and user feedback score, respectively; Q i 、C i 、E i 、U i Corresponding to the video generation model M i The scores on each evaluation indicator.
[0132] In the implementation of this application, the generation module 300 is further used to: perform quantitative processing on the evaluation data to obtain scores of multiple evaluation indicators.
[0133] In the embodiment of the present application, the device 10 of the embodiment of the present application further includes: a reminder module.
[0134] Among them, the reminder module is used to remind the user to adjust the current generation task or update the version of each video generation model if the comprehensive score of each video generation model is lower than the preset score before determining the target video generation model based on the comprehensive score.
[0135] In the embodiment of the present application, the device 10 of the embodiment of the present application further includes: an update module.
[0136] Among them, the update module is used to obtain the current evaluation data of the current generation task; and use the current evaluation data to update the evaluation data of the target video generation model to perform the historical video generation task.
[0137] It should be noted that the above explanation of the embodiment of the method for generating perceptual model training data is also applicable to the device for generating perceptual model training data of this embodiment, and will not be repeated here.
[0138] According to the device for generating perception model training data proposed in the embodiment of the present application, the element features and multiple evaluation indicators of the user's current generation task can be extracted, and the comprehensive scores of multiple video generation models can be calculated based on the multiple evaluation indicators and the evaluation data of historical generation tasks performed by multiple video generation models. Then, the most suitable target video generation model is matched according to the comprehensive score, and the target video generation model is used to generate training data of the perception model that matches the current generation task based on the required features, thereby realizing efficient generation of diversified and highly realistic driving scene videos through the video generation model, and using the driving scene videos as training data for the perception model, thereby improving the coverage and richness of driving scenes in the vehicle perception model training data set, and reducing the cost of using real collected data to train the vehicle-side perception model.
[0139] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0140] A memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .
[0141] When the processor 602 executes the program, the method for generating perception model training data provided in the above embodiment is implemented.
[0142] Furthermore, the electronic device further includes:
[0143] The communication interface 603 is used for communication between the memory 601 and the processor 602 .
[0144] The memory 601 is used to store computer programs that can be run on the processor 602 .
[0145] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0146] If the memory 601, processor 602, and communication interface 603 are implemented independently, the communication interface 603, memory 601, and processor 602 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0147] Optionally, in a specific implementation, if the memory 601, the processor 602 and the communication interface 603 are integrated on a chip, the memory 601, the processor 602 and the communication interface 603 can communicate with each other through an internal interface.
[0148] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0149] An embodiment of the present application also provides a vehicle, on which a perception model is provided, wherein the perception model is trained using training data generated by the above-mentioned method for generating perception model training data.
[0150] An embodiment of the present application also provides a computer-readable storage medium on which a computer program or instruction is stored. When the computer program or instruction is executed by a processor, the method for generating perception model training data as described above is implemented.
[0151] An embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed, implements the above-mentioned method for generating perception model training data.
[0152] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0153] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0154] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0155] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.
[0156] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
Claims
1. A method for generating perceptual model training data, characterized in that: The following steps are involved: Obtaining user evaluation data on the current generation task of the perception model training data and the historical generation tasks performed by multiple video generation models; Extracting required features and multiple evaluation indicators in the current generation task, wherein the required features include at least one of video theme, video length, style preference, image quality requirements, and specific elements, and the evaluation indicators include at least one of video quality, content matching, generation efficiency, and user feedback score; The comprehensive scores of the multiple video generation models are calculated according to the multiple evaluation indicators and the evaluation data, and a target video generation model is determined based on the comprehensive scores. The target video generation model generates training data for the perception model based on the required features.
2. The method for generating perceptual model training data according to claim 1, characterized in that: The step of calculating the comprehensive scores of the multiple video generation models based on the multiple evaluation indicators and the evaluation data further includes: Obtaining historical priorities of multiple evaluation indicators in the history generation task; Obtaining the current priorities of multiple evaluation indicators in the current generation task; Determining weights of multiple evaluation indicators in the current generation task according to at least one of the current priority and the historical priority; Extracting scores of multiple evaluation indicators from the evaluation data; The comprehensive score of each video generation model is calculated based on the weights and scores of the multiple evaluation indicators.
3. The method for generating perceptual model training data according to claim 1, characterized in that: The calculation formula of the comprehensive score is: S i =ω1×Q i +ω2×C i +ω3×E i +ω4×U i ; Among them, S i is the comprehensive score of the video generation model, ω1, ω2, ω3, and ω4 are the weights of video quality, content matching, generation efficiency, and user feedback score, respectively; Q i 、C i 、E i 、U i Corresponding to the video generation model M i The scores on each evaluation indicator.
4. The method for generating perceptual model training data according to claim 1, wherein: The extracting scores of multiple evaluation indicators from the evaluation data includes: The evaluation data is quantified to obtain scores of the multiple evaluation indicators.
5. The method for generating perceptual model training data according to claim 1, characterized in that: Before determining the target video generation model based on the comprehensive score, the method further includes: If the comprehensive score of each video generation model is lower than the preset score, the user is reminded to adjust the current generation task or update the version of each video generation model.
6. The method for generating perceptual model training data according to claim 1, characterized in that: Also includes: Obtaining current evaluation data of the currently generated task; The current evaluation data is used to update the evaluation data of the target video generation model for performing the historical video generation task.
7. A device for generating perception model training data, characterized in that: include: An acquisition module is used to obtain user evaluation data on the current generation task of the perception model training data and the historical generation tasks performed by multiple video generation models; an extraction module, configured to extract required features and multiple evaluation indicators in the current generation task, wherein the required features include at least one of video theme, video length, style preference, image quality requirements, and specific elements, and the evaluation indicators include at least one of video quality, content matching, generation efficiency, and user feedback score; A generation module is used to calculate the comprehensive scores of the multiple video generation models based on the multiple evaluation indicators and the evaluation data, determine the target video generation model based on the comprehensive scores, and the target video generation model generates training data for the perception model based on the required features.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for generating perceptual model training data as described in any one of claims 1 to 6.
9. A vehicle, characterized in that: The vehicle is provided with a perception model, wherein the perception model is trained using training data generated by the method for generating perception model training data according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: The computer program or instruction is executed by a processor to implement the method for generating perceptual model training data as described in any one of claims 1 to 6.