A key test case generalization method for automatic driving visual target detection

CN118506133BActive Publication Date: 2026-09-15SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410488861.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2026-09-15
Estimated Expiration
2044-04-23

AI Technical Summary

Technical Problem

[0005]现有技术的方法,无论是道路测试还是模拟器生成,均缺乏对关键环境因素的解析,因此,在获取测试图像时都存在相当程度的盲目性,且危险测试用例的比例较低;现有的图像合成方法,如神经迁移网络或生成对抗网络,通常需要大量的计算资源进行训练,这大大降低了测试效率和灵活性,同时,生成对抗网络还存在“模式崩塌”缺陷,经常导致所生成的测试图像高度雷同,缺乏多样性

Benefits of technology

[0037] In the key test case generalization method for autonomous driving visual target detection of the present invention, by collecting observational datasets and applying causal structure recognition methods, the key factors affecting the accuracy of the visual algorithm are first identified from the high-dimensional environment; furthermore, the optimal combination of various treatment levels of the key factors is obtained based on multi-objective search; the descriptions of the searched key factors are used as prompt words, and the test images are generalized in a targeted, efficient and low-cost manner by combining a diffusion model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118506133B_ABST
    Figure CN118506133B_ABST
Patent Text Reader

Abstract

The application provides a kind of automatic driving visual target detection key test case generalization method, comprising: obtaining observation image data under different environmental conditions;Observation image data is preprocessed, and the observation image data after preprocessing is collected into observation data set;Key environmental factors affecting target detection accuracy in observation data set are identified by causal reasoning;The challenge index of the preset test image and the coverage index of the measured algorithm are used to construct a multi-objective optimization problem;The best combination of key environmental factors is determined based on the multi-objective optimization problem;Fine-tune the open source diffusion model to obtain a standard diffusion model;The best combination of key environmental factors is used as a prompt word in combination with the standard diffusion model, and test cases are generalized in bulk through text-to-image and image-to-image methods.The method of the application uses the description of the searched key factors as a prompt word, and combines the diffusion model to realize the targeted, efficient and low-cost generalization of test images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and in particular to a method for generalizing key test cases for visual target detection in autonomous driving. Background Technology

[0002] Thanks to the significant advancements in artificial intelligence algorithms, particularly deep neural networks (DNNs), the market penetration rate of autonomous vehicles is gradually increasing. Key visual object detection algorithms (such as YOLO, SSD, and Faster-RCNN) are typically trained on data collected under specific conditions. However, due to the complex and harsh environmental challenges of the real world, existing algorithms often lead to erroneous perception results, significantly reducing the safety of autonomous vehicles.

[0003] To fully validate the robustness of visual perception algorithms, collecting test data under various challenging environments is crucial. Currently, the industry still heavily relies on road testing, which is costly and inefficient. Other researchers have attempted to obtain test images through high-fidelity simulators or image synthesis methods; however, due to a lack of understanding of challenging environments, these methods still cannot effectively balance the danger and diversity of test images.

[0004] Current advanced low-cost test image synthesis methods include affine transformation, generative adversarial networks (GANs), and neural style transfer networks. Generally, datasets collected on sunny days are relatively easy to acquire, but data from extreme weather conditions, such as rain, fog, or nighttime, are extremely rare. These extreme environments pose a significant challenge to current visual object detection algorithms for autonomous vehicles. Another typical example is image style transfer, including neural style transfer methods and generative adversarial networks. Neural style transfer transfers image content and style information between different layers in a convolutional neural network, generating new images with the target style by minimizing content and style losses.

[0005] Existing methods, whether for road testing or simulator generation, lack analysis of key environmental factors. As a result, there is a considerable degree of blindness in acquiring test images, and the proportion of dangerous test cases is low. Existing image synthesis methods, such as neural transfer networks or generative adversarial networks, usually require a large amount of computational resources for training, which greatly reduces testing efficiency and flexibility. At the same time, generative adversarial networks also suffer from the "pattern collapse" defect, which often leads to highly similar test images and a lack of diversity. Summary of the Invention

[0006] In view of this, the present invention provides a method for generalizing key test cases for visual target detection in autonomous driving, in order to solve the above problems.

[0007] This invention provides a method for generalizing key test cases for visual target detection in autonomous driving, comprising: acquiring observation image data under different environmental conditions; preprocessing the observation image data to form an observational dataset; identifying key environmental factors affecting target detection accuracy in the observational dataset through causal reasoning; constructing a multi-objective optimization problem by pre-setting challenge indicators for test images and coverage indicators for the algorithm under test; determining the optimal combination of the key environmental factors based on the multi-objective optimization problem; fine-tuning an open-source diffusion model to obtain a standard diffusion model; and combining the standard diffusion model with the optimal combination of the key environmental factors as prompt words to batch generalize test cases through text-to-image and image-to-image methods.

[0008] In another implementation of the present invention, the observed dataset needs to satisfy the negligibility assumption and the positive value assumption. The negligibility assumption requires that:

[0009]

[0010] The positive value assumption requires:

[0011]

[0012] in, Y represents the environmental conditions of interest, and Y represents the target detection accuracy. Denotes covariates, that is, in addition to Other background variables besides Y.

[0013] In another implementation of the present invention, the multi-objective optimization problem is expressed as:

[0014]

[0015] Where, vector v * The size is equal to the number of key factors screened through causal inference. It is a constrained space of integer encodings of all key factors. Given two scoring functions, expressed as:

[0016]

[0017] Wherein, s1(·) is the scoring function of the challenge indicator, and s2(·) is the scoring function of the coverage indicator of the algorithm under test.

[0018] In another implementation of the present invention, the scoring function s1(·) of the challenge indicator is expressed as:

[0019]

[0020] Where, N i and N j s1 represents the number of bounding boxes detected by algorithms i and j, respectively. A smaller s1 value indicates greater inconsistency in detection and leads to the generation of a more challenging test case dataset.

[0021] IOU k The calculation formula is as follows:

[0022] IOU k =Area(box) i ∩box j ) / Area(box i ∪box j )

[0023] Where Area is the pixel region of the detection box;

[0024]

[0025] Among them, Class k Used to determine whether the object categories of different algorithms are the same, Conf k This represents the detection confidence of the k-th bounding box;

[0026]

[0027] Here, "score" represents the confidence score.

[0028] In another implementation of the present invention, the scoring function s2(·) of the coverage index of the tested algorithm is expressed as:

[0029]

[0030] Where C represents the number of candidate algorithms being tested simultaneously;

[0031]

[0032] KMNC quantified the test image set. The range of values ​​for the lower neuron [low] o high o The degree of coverage, high o and low o These represent the upper and lower bounds of the neuron's activation value, respectively. o high oThe part is divided into m equal parts. Let... For the k-th part, φ(x,o) is the output value of neuron o∈O in the algorithm under test, where O is the set of neurons in the DNN, and for the test input... if Then the corresponding part It was considered a successful coverage.

[0033] In another implementation of the present invention, the text-to-image method guides the generation of key scenes based on the decoded prompts; the image-to-image method simultaneously selects the original image and prompts as conditions to transform the environment style of the open-source training set while preserving the image content.

[0034] In another aspect, the present invention provides a key test case generalization device for visual target detection in autonomous driving, comprising: a dataset acquisition module for acquiring observation image data under different environmental conditions; performing data preprocessing on the observation image data and aggregating the preprocessed observation image data into an observation dataset; a data processing module for identifying key environmental factors affecting target detection accuracy in the observation dataset through causal reasoning; constructing a multi-objective optimization problem by using preset challenge indicators of test images and coverage indicators of the tested algorithm; determining the optimal combination of the key environmental factors based on the multi-objective optimization problem; and a test case generalization module for fine-tuning an open-source diffusion model to obtain a standard diffusion model; and, combining the standard diffusion model with the optimal combination of the key environmental factors as prompt words, batch generalizing test cases through text-to-image and image-to-image methods.

[0035] In another aspect, the present invention provides an electronic device, characterized in that it includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for generalizing key test cases for autonomous driving visual target detection.

[0036] In another aspect, the present invention provides a computer storage medium, characterized in that a computer program is stored on the computer storage medium, and when the computer program is executed by a processor, it implements the steps in the above-described method for generalizing key test cases for autonomous driving visual target detection.

[0037] In the key test case generalization method for autonomous driving visual target detection of the present invention, by collecting observational datasets and applying causal structure recognition methods, the key factors affecting the accuracy of the visual algorithm are first identified from the high-dimensional environment; furthermore, the optimal combination of various treatment levels of the key factors is obtained based on multi-objective search; the descriptions of the searched key factors are used as prompt words, and the test images are generalized in a targeted, efficient and low-cost manner by combining a diffusion model. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. By reading the detailed description of the embodiments below, the advantages and benefits of the solutions will become clear to those skilled in the art. The accompanying drawings are only for illustrating preferred embodiments and are not intended to limit the present invention.

[0039] In the attached diagram:

[0040] Figure 1 This is a schematic diagram of a key test case generalization method for visual target detection in autonomous driving, according to an embodiment of the present invention.

[0041] Figure 2 This is a schematic diagram illustrating the generalization and application process of key test cases in one embodiment of the present invention.

[0042] Figure 3 This is a schematic diagram illustrating text annotation of images using a large language model, according to an embodiment of the present invention.

[0043] Figure 4 This is a schematic diagram of the main process for screening key environmental factors based on causal reasoning, according to an embodiment of the present invention.

[0044] Figure 5 This is a schematic diagram of a prompt word search based on multi-objective optimization, according to an embodiment of the present invention.

[0045] Figure 6 This is a schematic diagram of the Pareto front of feasible solutions for prompt words obtained through multi-objective search according to an embodiment of the present invention.

[0046] Figure 7 This is a schematic diagram illustrating the use of a finely tuned stable diffusion model to generalize test images in a "texture graph" manner, according to one embodiment of the present invention.

[0047] Figure 8(a) is a schematic diagram of the average detection accuracy of the target detection algorithm of an embodiment of the present invention on a variety of validation sets.

[0048] Figure 8(b) is a schematic diagram of the KMNC neuron coverage of an object detection algorithm according to an embodiment of the present invention on multiple validation sets.

[0049] Figure 9 This is a schematic diagram illustrating how fine-tuning a model in a generalized test scenario improves model performance, as shown in one embodiment of the present invention. Detailed Implementation

[0050] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.

[0051] Figure 1 This is a flowchart illustrating a key test case generalization method for visual target detection in autonomous driving, as provided in an embodiment of the present invention. Figure 1 As shown, this embodiment mainly includes the following steps:

[0052] Acquire observational image data under different environmental conditions.

[0053] The observed image data is preprocessed, and the preprocessed observed image data is then compiled into an observational dataset.

[0054] Key environmental factors affecting target detection accuracy in observational datasets are identified through causal reasoning.

[0055] A multi-objective optimization problem is constructed by setting a challenge metric for the test image and a coverage metric for the algorithm under test.

[0056] Determining the optimal combination of key environmental factors based on multi-objective optimization problems.

[0057] By fine-tuning the open-source diffusion model, a standard diffusion model is obtained.

[0058] For example, open-source diffusion models cannot recognize all custom keyword prompts, so fine-tuning is required to adapt to the style of the observed data, such as training a low-rank adaptive (LoRA) model so that the diffusion model can recognize the given text descriptions.

[0059] Preferably, the open-source diffusion model can be efficiently fine-tuned using LoRA (Low-Rank Adaptation). The core concept of LoRA is to decompose the gradient matrix (ΔW) of the cross-attention layer into a low-rank matrix during the training of the diffusion model.

[0060] ΔW=AB T

[0061] in, And d is much smaller than n.

[0062] Furthermore, A and B can be fine-tuned directly instead of W, which significantly reduces the required GPU computing resources. Subsequently, large language models (such as chatGPT) are used to perform text annotation on the collected CARLA images, such as... Figure 3 As shown, the data preparation process required for training the LoRA model is summarized.

[0063] It should be understood that the causal structure between variables is usually represented by a directed acyclic graph:

[0064]

[0065] Where E represents the set of causal relationships between variables, for example, express It has a direct causal effect on Y.

[0066] Representative causal structure learning methods aim to find the optimal graph structure g in the search space Ω. * :

[0067]

[0068] For example, the classic Greedy Equivalence Search (GES) starts from an empty graph and converges to a score metric (such as the Bayesian Information Criterion, BIC) by iteratively adding or removing edges.

[0069] This invention utilizes the Fast Greedy Equivalence Search (FGES) algorithm, a parallelized and optimized version of GES; this invention combines this algorithm with domain knowledge, such as... Figure 4 As shown, a causal graph of the visual object detection algorithm was constructed from randomly sampled observations of the CARLA dataset to further identify key environmental factors.

[0070] Combining the standard diffusion model, the optimal combination of key environmental factors is used as cue words, and test cases are generalized in batches through text-to-graph and graph-to-graph methods.

[0071] For example, such as Figure 2 As shown, generalized key test cases can be used as a supplementary set for training the algorithm model, so as to continuously improve the performance of the algorithm model in challenging environments.

[0072] In the key test case generalization method for autonomous driving visual target detection of the present invention, by collecting observational datasets and applying causal structure recognition methods, the key factors affecting the accuracy of the visual algorithm are first identified from the high-dimensional environment; furthermore, the optimal combination of various treatment levels of the key factors is obtained based on multi-objective search; the descriptions of the searched key factors are used as prompt words, and the test images are generalized in a targeted, efficient and low-cost manner by combining a diffusion model.

[0073] In another implementation of the present invention, the observation dataset needs to satisfy the negligibility assumption and the positive value assumption. The negligibility assumption requires:

[0074]

[0075] The positivity assumption requires:

[0076]

[0077] in, Y represents the environmental conditions of interest, and Y represents the target detection accuracy. Denotes covariates, that is, in addition to Other background variables besides Y.

[0078] For example, observational datasets are crucial for causal inference and fine-tuning diffusion models. Traditional causal inference methods rely on randomized controlled trials, which are either too expensive or impractical.

[0079] In this invention, the CARLA simulator can be used as an alternative method for quickly collecting images under different weather conditions. It is worth noting that observational datasets usually need to satisfy the negligibility assumption and the positive value assumption.

[0080] Specifically, the negligibility assumption requires:

[0081]

[0082] The positivity assumption requires:

[0083]

[0084] in, Y represents the environmental conditions of interest (e.g., fog, rain, time, etc.), and Y represents the target detection accuracy (e.g., mean accuracy, mAP). Denotes covariates, that is, in addition to Other background variables besides Y.

[0085] To satisfy these assumptions, images were collected using the principle of combinatorial testing. After perceptual hashing for deduplication, their no-reference image quality assessment metrics and detection accuracy scores were further extracted. These results were then compiled into a tabular observation dataset.

[0086] This invention introduces a causal inference method for the first time. By constructing an observational dataset and applying causal structure discovery methods, it identifies and mines key environmental factors affecting the accuracy of vision algorithms for autonomous vehicles from a high-dimensional environment. The innovation of this step lies in accurately identifying factors affecting algorithm performance in complex environments, which serve as the basis for generating subsequent test cases.

[0087] In another implementation of the present invention, the multi-objective optimization problem is expressed as:

[0088]

[0089] Where, vector v * The size is equal to the number of key factors screened through causal inference. It is a constrained space of integer encodings of all key factors. Given two scoring functions, expressed as:

[0090]

[0091] Where s1(·) is the scoring function for the challenge metric, and s2(·) is the scoring function for the coverage metric of the algorithm under test.

[0092] For example, in the process of image generation using the diffusion model in this invention, the description of the key environmental factors selected plays a crucial role. The prompt words can effectively guide the diffusion model and promote the generation of test images that are consistent with the expected prompts.

[0093] The identified key factors still exist in a rough description form. In practical applications, they need to be further subdivided at different processing levels. The combination of prompt words should be formalized into a multi-objective optimization problem to balance the risk level of the generated image and the coverage of the algorithm under test.

[0094] Specifically, given an n-dimensional vector representing the code of the cue word:

[0095] v={v time ,v fog ,…,v rainy}

[0096] The goal of this invention is to find an optimal vector v * The test image set can be generalized using a fine-tuned LoRA and an open-source diffusion model (ε).

[0097] v * Satisfy the following objective function:

[0098]

[0099] in, Given two scoring functions:

[0100]

[0101] vector v *The size is equal to the number of key factors screened through causal inference; It is a constrained space of integer codes for all key factors. For example, the "fog density" factor includes five levels: "no fog," "light fog," "moderate fog," "dense fog," and "extreme fog." fog The possible values ​​are 0, 1, 2, 3, and 4.

[0102] By introducing a multi-objective optimization problem, this invention proposes a method for searching the optimal combination of prompt words to balance the risk level of the generated test images and the coverage of the algorithm under test. The innovation lies in guiding the diffusion model to generate challenging test images that stimulate the robustness of the algorithm through systematic prompt word selection.

[0103] In another implementation of the present invention, the scoring function s1(·) of the challenge indicator is expressed as:

[0104]

[0105] Where, N i and N j s1 represents the number of bounding boxes detected by algorithms i and j, respectively. A smaller s1 value indicates greater inconsistency in detection and leads to the generation of a more challenging test case dataset.

[0106] IOU k The calculation formula is as follows:

[0107] IOU k =Area(box) i ∩box j ) / Area(box i ∪box j )

[0108] Where Area is the pixel region of the detection box;

[0109]

[0110] Among them, Class k Used to determine whether the object categories of different algorithms are the same, Conf k This represents the detection confidence of the k-th bounding box;

[0111]

[0112] Here, "score" represents the confidence score.

[0113] For example, considering that it is difficult to manually determine the danger of the generated images during the automatic generalization process, the images that induce inconsistent prediction results from multiple target detection algorithms are regarded as key test cases. In other words, key test cases will inevitably lead to at least one false detection or missed detection by the tested system. An evaluation function s1(·) is designed.

[0114] In another implementation of the present invention, the scoring function s2(·) of the coverage index of the tested algorithm is expressed as:

[0115]

[0116] Where C represents the number of candidate algorithms being tested simultaneously;

[0117]

[0118] KMNC quantified the test image set. The range of values ​​for the lower neuron [low] o high o The degree of coverage, high o and low o These represent the upper and lower bounds of the neuron's activation value, respectively. o high o The part is divided into m equal parts. Let... For the k-th part, φ(x,o) is the output value of neuron o∈O in the algorithm under test, where O is the set of neurons in the DNN, and for the test input... if Then the corresponding part It was considered a successful coverage.

[0119] For example, test cases generated using prompt words should not only be dangerous, but should also maximize the coverage of the algorithm being tested. To this end, a metric called K-multi-segment neuron coverage (KMNC) is introduced, and an evaluation function s2(·) is designed.

[0120] In another implementation of the present invention, the text-to-image method guides the generation of key scenes based on the decoded prompts; the image-to-image method selects both the original image and the prompts as conditions to transform the environment style of the open-source training set while preserving the image content.

[0121] For example, a reasonable approach to solving multi-objective optimization problems is to obtain a set of feasible solutions, each of which satisfies the optimization objective at an acceptable level without being dominated by any other solution (i.e., a Pareto optimal solution). To this end, the Fast Non-Dominated Sorting Genetic Algorithm (NSGA-II) is applied to obtain the optimal combination of prompt words, such as... Figure 5 As shown.

[0122] Preferably, once the optimal combination of prompt words is obtained through multi-object search, two methods can be used to generalize key test cases: "text-to-image" and "image-to-image". The former relies solely on the decoded text prompts to guide the generation of key scenes; the latter selects both the original image and the prompt words as conditions to transform the environmental style of the initial training set while preserving the image content. The robustness of the object detection algorithm can be further analyzed.

[0123] This invention proposes two generalization methods: the first is the "text-to-image" stage, which relies solely on decoded text prompts to guide the generation of key scenes; the second is the "image-to-image" stage, which transforms the environmental style of the initial training set by selecting a combination of original images and prompt words. This allows for a more comprehensive evaluation of the visual algorithm performance of autonomous vehicles. Compared to a single generalization method, this strategy of combining two approaches makes the test cases more extensive and complex. The innovation lies in using two different methods to generate test cases to comprehensively evaluate the algorithm's performance.

[0124] A non-dominated sorting genetic algorithm (NSGA-II) is introduced for prompt word search. This algorithm improves the quality of prompt word combinations by obtaining a set of feasible solutions, ensuring that each solution satisfies the multi-objective optimization objective at an acceptable level. The innovation lies in using an effective evolutionary algorithm to search for the optimal prompt word combination.

[0125] In another implementation of the present invention, the feasibility has been fully demonstrated through extensive experiments, specifically including:

[0126] (1) Identify the optimal combination of prompt words

[0127] Thirteen variables were selected from CARLA, generating a total of 4513 combinations of environmental conditions. These factors included various weather parameters (rain, fog, sunset, etc.), different types of traffic participants (cars, trucks, bicycles, pedestrians, etc.), and various driving scenarios such as cities, highways, and towns. A no-reference image quality metric was applied to each weather condition. Subsequently, the collected data was tested using an object detection algorithm pre-trained on the SHIFT dataset (an open-source benchmark dataset for CARLA) to ensure benchmark performance. Following the method outlined in this invention, a causal graph was successfully constructed, from which five key environmental factors were selected.

[0128] Meanwhile, the multi-objective optimization NSGA-II algorithm was configured to perform a maximum of 20 evolutionary iterations and a population size of 100. Each possible solution was decoded as a hint, and each possible solution in the search process generated a test set of 50 images through a diffusion model. The result was a Pareto optimal set consisting of 15 feasible solutions, and the Pareto front was as follows: Figure 6 As shown.

[0129] like Figure 6 As shown, a feasible solution represents a solution that satisfies all constraints within the framework of a multi-objective problem. In this context, a feasible solution means an encoded vector of a set of cue words, which, when decoded, produces a text description. For example, one feasible solution v = [5 3 3 2 2], after decoding, represents: "afternoon", "moderate fog", "highway", "light rain", "dense traffic flow".

[0130] (2) Verify test cases

[0131] To assess the challenge posed by the generalized test images of this invention, a fine-tuned diffusion model collected 40 images for each candidate prompt word combination in a "text-to-image" manner, generating a validation set containing approximately 600 images. Figure 7 This invention demonstrates some of the high-quality test images generated by the present invention.

[0132] In addition, a validation dataset containing 792 images was constructed using the "image-to-image" method. This process preserved the content of the original training images and adopted the style described by the cue words to transform the original training dataset into various challenging environments. The accuracy and neuron coverage of the tested algorithms on various datasets are shown in Figure 8.

[0133] As shown in Figure 8(a), the object detection algorithms trained on the open-source SHIFT dataset exhibited performance degradation on both randomly collected observation datasets and the validation set constructed in this invention. Specifically, the "Text2Scene" dataset was the most challenging for evaluating the algorithms, with mAP@50 for YOLO, SSD, and Faster-RCNN dropping to 0.351, 0.212, and 0.291, respectively. The object detectors also performed poorly on the "Image2Scene" dataset, with an average detection accuracy of only 0.37, significantly lower than the 0.537 on the observed CARLA dataset and 0.81 on the SHIFT dataset.

[0134] Experimental results demonstrate that the proposed method effectively identifies challenging key environmental conditions and, combined with the few-shot generation capability of the diffusion model, efficiently generalizes key test cases. It should also be noted that the object detection algorithm achieves improved accuracy on the "Image2Scene" dataset compared to the "Text2Scene" dataset. This is because "Image2Scene" relies not only on text prompts but also considers the content of the original training images. While this method limits the degrees of freedom in generation, it also makes the generated driving scenes closer to the distribution of the original SHIFT training data, thereby improving the model's detection performance.

[0135] Furthermore, the statistical results in Figure 8(b) demonstrate that the generalized test set not only generates challenging test cases but also takes into account neuron coverage. Specifically, the evaluated YOLO algorithm achieves neuron coverage of 0.53 and 0.59 on the “Text2Scene” and “Image2Scene” datasets, respectively, which is close to the neuron coverage of the original SHIFT validation set (0.63).

[0136] (3) Model reinforcement

[0137] This invention further explores the possibility of improving the performance of the tested algorithm by increasing the generalization of test scenarios under harsh environmental conditions. Specifically, this invention reintegrates the generalization key test cases from "Text2Scene" and "Image2Scene" into the initial training set. Subsequently, the object detection algorithm is fine-tuned, and its performance on all 4513 environment combinations on the observational CARLA dataset is analyzed. The results are as follows: Figure 9 As shown.

[0138] Figure 9 This indicates that the performance of the fine-tuned algorithm was improved to some extent on the observational CARLA dataset with the largest coverage. Specifically, the median detection accuracy of YOLO, SSD, and Faster-RCNN improved by 4.9%, 8.13%, and 6.8%, respectively. In other words, the method of generalizing key test cases proposed in this invention effectively promotes the continuous improvement of the algorithm model. As the types and number of generalized test cases increase, the model performance will be further improved.

[0139] This invention addresses the shortcomings of existing visual target detection methods, such as the lack of key environmental factors and low efficiency in generating test cases. It proposes a visual test case generalization method based on causal inference and diffusion models to specifically enhance the generation of key test images, thereby effectively improving the efficiency of safety assessment for autonomous vehicles.

[0140] In another aspect, the present invention provides a key test case generalization apparatus for visual object detection in autonomous driving, comprising:

[0141] Dataset acquisition module: Acquires observation image data under different environmental conditions; performs data preprocessing on the observation image data; and compiles the preprocessed observation image data into an observational dataset.

[0142] Data processing module: Identifies key environmental factors affecting target detection accuracy in observational datasets through causal reasoning; constructs a multi-objective optimization problem using pre-set challenge indicators of test images and coverage indicators of the tested algorithm; and determines the optimal combination of key environmental factors based on the multi-objective optimization problem.

[0143] Test case generalization module: Fine-tunes the open-source diffusion model to obtain a standard diffusion model; Combined with the standard diffusion model, it uses the optimal combination of key environmental factors as prompt words to generalize test cases in batches through text-to-graph and graph-to-graph methods.

[0144] In the key test case generalization system for autonomous driving visual target detection of the present invention, by collecting observational datasets and applying causal structure recognition methods, the key factors affecting the accuracy of the visual algorithm are first identified from the high-dimensional environment; furthermore, the optimal combination of various treatment levels of the key factors is obtained based on multi-objective search; the descriptions of the searched key factors are used as prompt words, and the test images are generalized in a targeted, efficient and low-cost manner by combining a diffusion model.

[0145] In another aspect of the present invention, the electronic device includes: a processor, a memory, and a communication bus and a communication interface.

[0146] in:

[0147] The processor, memory, and communication interface communicate with each other via a communication bus.

[0148] A communication interface is used to communicate with other electronic devices or servers.

[0149] The processor is used to execute programs, specifically the steps of any of the key test case generalization methods for autonomous driving visual object detection in the above embodiments.

[0150] Specifically, the program may include program code, which includes computer operation instructions.

[0151] The processor may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0152] Memory is used to store programs. Memory may include high-speed RAM, and may also include non-volatile memory, such as at least one disk drive.

[0153] Specifically, the program can be used to cause the processor to execute steps to implement any of the key test case generalization methods for autonomous driving visual object detection described in the embodiments. The specific implementation of each step in the program can be found in the corresponding descriptions of the steps and units executed in any of the key test case generalization methods for autonomous driving visual object detection described above, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments.

[0154] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods of various embodiments of this application.

[0155] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0156] Specific embodiments of the invention have now been described. Other embodiments are within the scope of the appended claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0157] It should be noted that all directional indicators (such as up, down, left, right, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.

[0158] In the description of this invention, the terms "first" and "second" are used only for convenience in describing different components or names, and should not be construed as indicating or implying a sequential relationship, relative importance, or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" and "second" may explicitly or implicitly include at least one of that feature.

[0159] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0160] It should be noted that although specific embodiments of the present invention have been described in detail with reference to the accompanying drawings, this should not be construed as limiting the scope of protection of the present invention. Various modifications and variations that can be made by those skilled in the art without inventive effort within the scope described in the claims still fall within the scope of protection of the present invention.

[0161] The examples of the embodiments of the present invention are intended to concisely illustrate the technical features of the embodiments of the present invention, so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and are not intended to be an improper limitation of the embodiments of the present invention.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generalizing key test cases for visual object detection in autonomous driving, characterized in that, include: Acquire observational image data under different environmental conditions; The observed image data is preprocessed, and the preprocessed observed image data is compiled into an observational dataset. Key environmental factors affecting target detection accuracy in the observed dataset are identified through causal reasoning. A multi-objective optimization problem is constructed by using a scoring function for the challenge metric of the pre-set test image and a scoring function for the coverage metric of the algorithm under test. Determine the optimal combination of the key environmental factors based on the multi-objective optimization problem; The open-source diffusion model was fine-tuned to obtain the standard diffusion model; Combining the standard diffusion model, the optimal combination of the key environmental factors is used as prompt words, and test cases are generalized in batches through text-to-graph and graph-to-graph methods. in: The scoring function for the challenge indicator Represented as: in, and Representing the algorithms respectively i and j The number of bounding boxes detected; The calculation formula is as follows: in, Area It is the pixel area of ​​the detection box; in, Used to determine whether the object categories of different algorithms are the same. Indicates the first k Detection confidence of each bounding box; Wherein, score represents the confidence score; The scoring function for the coverage metric of the algorithm under test. Represented as: in, This indicates the number of candidate algorithms being tested simultaneously. KMNC quantified the test image set. Lower neuron value range The degree of coverage, and These represent the upper and lower bounds of the neuron's activation value, respectively. Divided into An equal part For the first k Each part Neurons in the algorithm under test The output value, It is a collection of neurons in a DNN, for the test input ,if Then the corresponding part It was considered a successful coverage; The multi-objective optimization problem is expressed as: Where, vector The size is equal to the number of key environmental factors screened through causal inference. It is a constrained space of integer codes for all key environmental factors. Given two scoring functions, expressed as: in, The scoring function for the aforementioned challenge indicator. is the scoring function for the coverage index of the algorithm under test.

2. The method according to claim 1, characterized in that, The observed dataset needs to satisfy the negligibility and positivity assumptions. The negligibility assumption requires that: The positive value assumption requires: in, Represents the environmental conditions of interest. This represents the target detection accuracy. Denotes covariates, that is, in addition to and Other background variables besides.

3. The method according to claim 1, characterized in that, The text-to-image method guides the generation of key scenes based on decoded prompts; The image-to-image method selects both the original image and the prompt words as conditions to transform the environmental style of the initial training set while preserving the image content.

4. A key test case generalization device for visual target detection in autonomous driving, characterized in that, include: Dataset acquisition module: Acquires observational image data under different environmental conditions; The observed image data is preprocessed, and the preprocessed observed image data is compiled into an observational dataset. Data processing module: Identifies key environmental factors affecting target detection accuracy in the observed dataset through causal reasoning; constructs a multi-objective optimization problem using a scoring function of a challenge index for the test image and a scoring function of a coverage index for the tested algorithm; and determines the optimal combination of the key environmental factors based on the multi-objective optimization problem. Test case generalization module: Fine-tunes the open-source diffusion model to obtain a standard diffusion model; Combines the standard diffusion model with the optimal combination of the key environmental factors as prompt words, and generalizes test cases in batches through text-to-graph and graph-to-graph methods. in: The scoring function for the challenge indicator Represented as: in, and Representing the algorithms respectively i and j The number of bounding boxes detected; The calculation formula is as follows: in, Area It is the pixel area of ​​the detection box; in, Used to determine whether the object categories of different algorithms are the same. Indicates the first k Detection confidence of each bounding box; Where, score represents the confidence score. The scoring function for the coverage metric of the algorithm under test. Represented as: in, This indicates the number of candidate algorithms being tested simultaneously. KMNC quantified the test image set. Lower neuron value range The degree of coverage, and These represent the upper and lower bounds of the neuron's activation value, respectively. Divided into An equal part For the first k Each part Neurons in the algorithm under test The output value, It is a collection of neurons in a DNN, for the test input ,if Then the corresponding part It was considered a successful coverage; The multi-objective optimization problem is expressed as: Where, vector The size is equal to the number of key environmental factors screened through causal inference. It is a constrained space of integer codes for all key environmental factors. Given two scoring functions, expressed as: in, The scoring function for the aforementioned challenge indicator. is the scoring function for the coverage index of the algorithm under test.

5. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the key test case generalization method for autonomous driving visual target detection as described in claims 1-3.

6. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, implements the steps in the key test case generalization method for autonomous driving visual target detection as described in claims 1-3.