Vehicle-pedestrian collision working condition sample generation method and device, equipment and medium
Patent Information
- Application Number
- CN202610771300.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-06-01
AI Technical Summary
[0004]本申请的目的在于提供一种车辆-行人碰撞工况样本生成方法及装置、设备、介质,以解决现有技术存在的核心区样本密度不足、边界空心化、核心与全局过渡突兀、训练数据集覆盖不全面等问题
本申请提供的车辆-行人碰撞工况样本生成方法通过核心区内部区域+边带区域的分层设计与全局区多分区+软环抑制概率模型,在有限样本量下同时实现了法规核心工况的高密度覆盖和真实事故全域的无空洞覆盖,极大提升了机器学习模型在关键区域的拟合精度;软环抑制概率模型有效缓解了任意维度下的边界贴边聚团与掏空难题,使核心到全局的样本密度实现平滑过渡,边界附近行人伤害预测更可靠;专用极值点生成机制确保高风险工况全面覆盖,避免了机器学习训练的安全死角。
Smart Images

Figure CN122332965B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning, and more specifically, to a method, apparatus, device, and medium for generating vehicle-pedestrian collision condition samples. Background Technology
[0002] Vehicle-pedestrian collision testing is a core component of the development of Automatic Emergency Braking (AEB), pedestrian protection, and autonomous driving decision-making systems. The generated crash test samples require massive training data produced through multi-rigid-body dynamics simulations to enable machine learning models to learn and predict pedestrian injuries and optimal decision-making strategies. Existing methods primarily employ fixed crash conditions, global Monte Carlo sampling, or standard Latin hypercube sampling, which suffer from the following significant drawbacks: failure to distinguish between the regulatory core crash condition zone and real-world extreme crash conditions leads to insufficient sample density and hollowed-out boundaries in the core area, hindering accurate fitting of the nonlinear laws governing pedestrian injuries; samples tend to cluster or hollow out at the core area boundaries, resulting in abrupt transitions between the core and global areas and poor generalization ability of machine learning models near the boundaries; lack of dedicated generation mechanisms for high-risk combinations such as high-speed, steep slopes, and other high-risk scenarios, resulting in incomplete training dataset coverage; insufficient control over nearest neighbor distance and spatial coverage, leading to clustering or large voids, requiring a significant increase in sample size to meet training needs; poor adaptability to different vehicle models, resulting in high costs for manual parameter tuning. These shortcomings lead to low efficiency, high cost, and difficulty in guaranteeing the quality of machine learning training datasets. While existing patents cover collision scenario design, none of them offer systematic solutions for soft constraints, anti-hollowing, and smooth boundary transitions in the three-dimensional parameter space of the vehicle-pedestrian interaction.
[0003] In view of the above, this application is hereby submitted. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and medium for generating vehicle-pedestrian collision condition samples, in order to solve the problems of insufficient core area sample density, hollowed-out boundaries, abrupt transition between core and global areas, and incomplete coverage of training datasets in the existing technology.
[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, this application provides a method for generating vehicle-pedestrian collision scenario samples, including: Based on the preset total number of samples for each vehicle model, the number of samples in each region is determined. The samples in each region include core region samples, global region samples, and extreme value region samples. The core region includes an inner region and a side region. The global region includes a ring region and an outer region. The ring region is the region close to the core region, and the outer region is the region far away from the core region. The core area samples are generated based on the number of core area samples and the volume ratio of each sub-region in the core area; The number of samples in each sub-region of the global region is determined based on the number of samples in the global region and the volume ratio of each sub-region in the global region. The acceptance probability of ring candidate samples in the global region is calculated using a soft ring suppression probability model. If the acceptance probability exceeds a preset probability, the ring candidate sample is determined as an acceptable sample. The global region sample is generated based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region. The extreme value region sample is generated based on spatial corner points, surface center points, and high-risk combinations.
[0006] In some technical solutions, generating the core region samples based on the number of core region samples and the volume ratio of each sub-region in the core region includes: Based on the number of samples in the core area and the volume ratio of each sub-region in the core area, the number of samples in the internal region and the number of samples in the side zone region are determined respectively. The core area samples are generated based on the number of samples in the internal region and the number of samples in the side zone region.
[0007] In some technical solutions, before calculating the acceptance probability of ring candidate samples in the global region using a soft ring suppression probability model, a step of constructing a soft ring suppression probability model is included, including: A soft ring suppression probability model is constructed based on the minimum acceptance probability, the normalized distance from the annular candidate sample to the core region, the annular region width, and the suppression intensity index.
[0008] In some technical solutions, the soft loop suppression probability model is as follows: ; Where p is the acceptance probability of the annular candidate sample, p_min is the minimum acceptance probability, dist is the normalized distance from the annular candidate sample to the core region, ringBand is the width of the annular region, and α is the suppression intensity index.
[0009] In some technical solutions, generating the global region sample based on the sample quantity of each sub-region of the global region, the acceptable samples, and the parameter range of the core region includes: Based on the number of samples in each sub-region of the global region and the acceptable samples, a first global region sample is generated; Based on the parameter range of the core region, soft transition samples are generated, and all soft transition samples fall within the core region.
[0010] In some technical solutions, after generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes: Merge core region samples, global region samples, and extreme value region samples, and transform them into a normalized space; Based on the distance of each sample to the nearest other sample in the normalized space, the cluster point is determined. The cluster point is the sample point whose Euclidean distance to the nearest other sample point is less than a preset nearest neighbor distance threshold. For the cluster point, candidate samples are regenerated within the region to which the cluster point belongs, and the global external points are further screened and replaced using the soft loop suppression probability model until the distance constraint is met.
[0011] In some technical solutions, after generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes: Based on the generated samples from each region, the values of the evaluation indicators are calculated. The evaluation indicators include the minimum nearest neighbor distance for the entire region, the 95th percentile nearest neighbor distance, the 95th percentile of external spatial coverage, and the proportion of samples in the annular region. If the values of the evaluation indicators do not all meet the standards, a maximum number of retries is set, samples for each region are regenerated, and the result with the highest comprehensive score is selected as the final sample from the retries.
[0012] Secondly, this application provides a vehicle-pedestrian collision condition sample generation device, comprising: The module for determining the number of samples in each region is used to determine the number of samples in each region based on the preset total number of samples for the vehicle model. The samples in each region include core region samples, global region samples, and extreme value region samples. The core region includes an inner region and a side region. The global region includes a ring region and an outer region. The ring region is the region close to the core region, and the outer region is the region far away from the core region. The core area sample generation module is used to generate the core area sample based on the number of core area samples and the volume ratio of each sub-region in the core area; The global region sample generation module is used to determine the number of samples in each sub-region of the global region based on the number of global region samples and the volume ratio of each sub-region in the global region; calculate the acceptance probability of the ring candidate samples in the global region using a soft ring suppression probability model; and determine the ring candidate samples as acceptable samples if the acceptance probability exceeds a preset probability; and generate the global region samples based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region. The extreme value region sample generation module is used to generate the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations.
[0013] Thirdly, this application provides an electronic device, comprising: At least one processor, and a memory communicatively connected to at least one of the processors; The memory stores instructions that can be executed by at least one of the processors, which are executed by at least one of the processors to enable at least one of the processors to perform the method described above.
[0014] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to perform the above-described method.
[0015] Compared with the prior art, the beneficial effects of this application are as follows: The vehicle-pedestrian collision scenario sample generation method provided in this application achieves high-density coverage of regulatory core scenarios and void-free coverage of the entire real accident domain with a limited sample size through a hierarchical design of the core area + side area and a multi-partition global area + soft ring suppression probability model. This greatly improves the fitting accuracy of the machine learning model in key areas. The soft ring suppression probability model effectively alleviates the problem of boundary clumping and hollowing out in any dimension, enabling a smooth transition of sample density from the core to the global domain, and making pedestrian injury prediction near the boundary more reliable. The dedicated extreme point generation mechanism ensures comprehensive coverage of high-risk scenarios and avoids safety blind spots in machine learning training.
[0016] Furthermore, the local repair mechanism targeting only the clumping points (combining type constraints and soft loop suppression) results in a high generation success rate, eliminates the need for overall sampling restart, and significantly improves computational efficiency, making it particularly suitable for batch production of multi-vehicle and large-scale multi-rigid-body simulation datasets. It naturally supports expansion from three dimensions to any higher-dimensional parameter space (e.g., adding vehicle features, road surface adhesion coefficients, weather conditions, etc. in the future) without changing the core framework, only requiring adjustments to dimension D and the hierarchical strategy, greatly enhancing the method's versatility and foresight.
[0017] Furthermore, the multi-index closed-loop constraint and scoring optimization mechanism, along with the directly available Excel output and visual diagnostic charts, significantly reduces the data preparation cost and time for vehicle-pedestrian collision machine learning training, and improves the efficiency of model training and the reliability of the final safety performance assessment. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the method for generating vehicle-pedestrian collision scenario samples provided in this application.
[0020] Figure 2 This is a histogram comparison of vehicle speed, road slope, and collision angle parameters in the embodiments of this application (core area and global area, using tiledlayout 3×2 layout).
[0021] Figure 3 This is a vehicle speed-road slope plane distribution diagnostic map in the embodiments of this application, showing the core rectangle, the equidistant contour of the ring zone, and the sampling points colored by point type (orange is the core area sample, blue is the global external sample, cyan is the soft transition sample, and red is the extreme value sample).
[0022] Figure 4 It is the nearest neighbor distance histogram of the normalized three-dimensional space in the embodiments of this application.
[0023] Figure 5 This is a scatter matrix diagram of the normalized three-dimensional space in an embodiment of the present invention.
[0024] Figure 6 This is a schematic diagram of the vehicle-pedestrian collision condition sample generation device provided in this application.
[0025] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0026] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0027] Example 1 Figure 1 This is a flowchart of a method for generating vehicle-pedestrian collision condition samples provided in this embodiment. This method can be executed by a vehicle-pedestrian collision condition sample generation device, which can be composed of software and / or hardware and is generally integrated into an electronic device, such as a computer. For ease of understanding, each step in the generation method of this embodiment is executed by an electronic computer.
[0028] like Figure 1 As shown, this embodiment provides a method for generating vehicle-pedestrian collision scenario samples, including the following steps: S110. Based on the preset total number of samples for the vehicle model, determine the number of samples in each region. The samples in each region include core region samples, global region samples, and extreme value region samples. The core region includes an inner region and a side region. The global region includes a ring region and an outer region. The ring region is the region close to the core region, and the outer region is the region far away from the core region.
[0029] The term "close to the core area" refers to the fact that the distance from the global sample point to the boundary of the core area in the normalized vg plane is no greater than 0.025. The area that meets this condition is defined as the annular area, and the area with a distance greater than 0.025 is defined as the peripheral area.
[0030] Multiple vehicle models can be set, and each model generates a preset total number of samples N, which are divided into core area samples N_core, global area samples N_global (the global area is further divided into ring area and peripheral area), and extreme value area samples N_extreme; define the physical global range and core working condition range of D parameters; set the side area width, ring area width, minimum nearest neighbor distance threshold, external space coverage threshold, and optimize sampling parameters.
[0031] S120. Generate the core area samples based on the number of core area samples and the volume ratio of each sub-region in the core area.
[0032] Optionally, generating the core region samples based on the number of core region samples and the volume ratio of each sub-region in the core region includes: Based on the number of samples in the core area and the volume ratio of each sub-region in the core area, the number of samples in the internal region and the number of samples in the side zone region are determined respectively. The core area samples are generated based on the number of samples in the internal region and the number of samples in the side zone region.
[0033] The core region is a core hyperrectangle. During the generation of core region samples, multidimensional optimized Latin hypercube sampling can be used (achieved through "2+1", "3+1", or multi-stage pairing methods): prioritizing the uniformity of key two-dimensional or three-dimensional projection surfaces, and then pairing other dimensions step by step. The number of samples in the side regions is allocated according to the volume ratio of each sub-region. After sampling, these samples are merged with the internal samples to ensure dense samples near the core region boundary and avoid hollowing out.
[0034] S130. Determine the number of samples in each sub-region of the global region based on the number of samples in the global region and the volume ratio of each sub-region in the global region.
[0035] The global region is divided into multiple continuous sub-regions, and the number of samples is allocated according to the volume ratio of each sub-region. Multidimensional optimized Latin hypercube sampling is used in each sub-region.
[0036] S140. The acceptance probability of the ring candidate sample in the global region is calculated using the soft ring suppression probability model. If the acceptance probability exceeds the preset probability, the ring candidate sample is determined as an acceptable sample.
[0037] By employing a soft-ring suppression probability model to calculate the acceptance probability of ring candidate samples in the global region, and then selecting acceptable samples, global uniform coverage and natural connection with the core region can be achieved. The preset probability can be 0.25. This value, as the minimum acceptance probability of ring candidate samples, is empirically set based on sample spatial coverage requirements and ring region continuity requirements, and its value is within the recommended range of 0.20-0.35. This can suppress excessive sample aggregation near the core region boundary while avoiding obvious gaps in the ring region.
[0038] Optionally, before calculating the acceptance probability of ring candidate samples in the global region using the soft ring suppression probability model, the method further includes a step of constructing the soft ring suppression probability model, including: A soft ring suppression probability model is constructed based on the minimum acceptance probability, the normalized distance from the annular candidate sample to the core region, the annular region width, and the suppression intensity index.
[0039] Optionally, the soft loop suppression probability model is: ; Where p is the acceptance probability of the annular candidate sample, p_min is the minimum acceptance probability, dist is the normalized distance from the annular candidate sample to the core region, ringBand is the width of the annular region, and α is the suppression intensity index. α is used to adjust the rate at which the acceptance probability of the annular candidate sample increases with its distance from the core region. The larger α is, the lower the acceptance probability near the core region boundary, and the less likely the sample is to approach the core region boundary; the smaller α is, the faster the acceptance probability increases, and the weaker the suppression effect. In this embodiment, α is set to 2.2 in the generation stage and 1.8 in the repair stage. This value is an empirical adjustment parameter used to achieve a balance between excessive sample aggregation near the suppression boundary and avoiding the complete hollowing out of the annular region.
[0040] S150. Generate the global region sample based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region.
[0041] Optionally, the global region samples are generated based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region, including: Based on the number of samples in each sub-region of the global region and the acceptable samples, a first global region sample is generated; Based on the parameter range of the core region, soft transition samples are generated, and all soft transition samples fall within the core region.
[0042] In this embodiment, the global region samples include first global region samples and soft transition samples. Soft transition samples are few in number, generated within the core region, and all fall within the core region to further smooth the boundary. Forty soft transition samples are generated for each vehicle model. These soft transition samples are generated based on the parameter range of the core region. Specifically, under the condition that both velocity v and parameter g are within the core region and the angle parameter covers the entire angle range, an optimized Latin hypercube sampling method is used to generate soft transition samples to mitigate abrupt boundary changes between core region samples and global samples outside the core region.
[0043] S160. Generate the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations.
[0044] The system generates extreme sample points covering all corner points, surface center points, and several high-risk combinations in the coverage space, and applies small random perturbations to avoid complete overlap. Among them, the high-risk combination refers to the extreme parameter combination composed of high velocity, boundary g value, and large angle, which is used to enhance the coverage capability of DOE (Design of Experiments) samples for boundary conditions and potentially sensitive areas.
[0045] Optionally, after generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes: Merge core region samples, global region samples, and extreme value region samples, and transform them into a normalized space; Based on the distance of each sample to the nearest other sample in the normalized space, the cluster point is determined. The cluster point is the sample point whose Euclidean distance to the nearest other sample point is less than a preset nearest neighbor distance threshold. For the cluster point, candidate samples are regenerated within the region to which the cluster point belongs, and the global external points are further screened and replaced using the soft loop suppression probability model until the distance constraint is met.
[0046] A sample point is defined as a cluster point when its Euclidean distance to the nearest other sample point in the normalized 3D space S3=[v_scaled,g_scaled,asin_scaled] is less than a preset nearest neighbor distance threshold. In this embodiment, the preset nearest neighbor distance threshold is 0.010; when the Euclidean distance of a sample point to the nearest other sample point is less than 0.010, it is determined as a cluster point.
[0047] To eliminate the dimensional differences between different parameters, the physical parameters are mapped to [0,1]. A standard space is defined, and each sample is labeled with its type (core region, global region, extreme value region). For the D-dimensional mapping space, the normalized mapping model in this embodiment is: x i _scaled = (xi x i _min) / (x i _max x i (_min) (i=1,2,…,D), x i Let x be the original value of the i-th physical parameter. i _min is the lower bound of this parameter in the design space, x i _max is the upper bound of this parameter in the design space, x i _scaled represents the normalized standard coordinates, x i _scaled∈[0,1].
[0048] By calculating the distance from each sample to the nearest other sample in the normalized space, for "crowded points" with too small a distance, candidate samples are regenerated in their respective regions according to their type, and the soft loop suppression probability model is applied to the global external points for screening and replacement. The number of rounds and the number of samples in each round are set until the distance constraint is met, thereby achieving local repair optimization.
[0049] The patching process is not limited to samples in the outer region of the global region, but is carried out in two steps. The first step patches all clumping points in the samples, including core region samples, global outer samples, soft transition samples, and extreme value samples; when regenerating candidate samples, the generation area of the candidate samples is limited according to the type of the clumping point. The second step further patches the sample points outside the core region by applying distance constraints to the outer region. Therefore, "global outer points" refers to sample points located outside the core region, including external samples in the ring region and the outer region, not just samples in the outer region. In addition, other types of clumping points also need to be screened and replaced, but the screening method is different. For core region samples and soft transition samples, candidate samples are regenerated within the core region; for extreme value samples, candidate samples are regenerated throughout the entire region; for global outer points, candidate samples are regenerated outside the core region, and a soft ring suppression probability model is used for screening when the candidate samples are close to the core region boundary. In other words, all clumping points can be replaced; only global outer points are additionally screened using soft ring suppression.
[0050] Optionally, after generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes: Based on the generated samples from each region, the values of the evaluation indicators are calculated. The evaluation indicators include the minimum nearest neighbor distance for the entire region, the 95th percentile nearest neighbor distance, the 95th percentile of external spatial coverage, and the proportion of samples in the annular region. If the values of the evaluation indicators do not all meet the standards, a maximum number of retries is set, samples for each region are regenerated, and the result with the highest comprehensive score is selected as the final sample from the retries.
[0051] Among them, the external spatial coverage is used to determine whether the area outside the core region is "covered" by samples. The better the coverage, the more uniform the samples in the external region; the worse the coverage, the more likely some external regions are lacking samples.
[0052] Optionally, after generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes a parallel output step: supporting simultaneous parallel calculation of multiple vehicle models, ultimately restoring the physical parameters, outputting an Excel spreadsheet containing Chinese vehicle model names, and attaching a distribution diagnostic map.
[0053] The aforementioned vehicle-pedestrian collision scenario sample generation method achieves high-density coverage of core regulatory scenarios and void-free coverage of the entire real accident domain with a limited sample size through a hierarchical design of the core area plus the side area and a multi-partition global area plus a soft-loop suppression probability model. This greatly improves the fitting accuracy of the machine learning model in key areas. The soft-loop suppression probability model completely solves the problem of boundary clumping and hollowing out in any dimension, enabling a smooth transition of sample density from the core to the global domain, and making pedestrian injury prediction near the boundary more reliable. The dedicated extreme point generation mechanism ensures comprehensive coverage of high-risk scenarios and avoids safety blind spots in machine learning training.
[0054] Furthermore, the local repair mechanism targeting only the clumping points (combining type constraints and soft loop suppression) results in a high generation success rate, eliminates the need for overall sampling restart, and significantly improves computational efficiency, making it particularly suitable for batch production of multi-vehicle and large-scale multi-rigid-body simulation datasets. It naturally supports expansion from three dimensions to any higher-dimensional parameter space (e.g., adding vehicle features, road surface adhesion coefficients, weather conditions, etc. in the future) without changing the core framework, only requiring adjustments to dimension D and the hierarchical strategy, greatly enhancing the method's versatility and foresight.
[0055] Furthermore, the multi-index closed-loop constraint and scoring optimization mechanism, along with the directly available Excel output and visual diagnostic charts, significantly reduces the data preparation cost and time for vehicle-pedestrian collision machine learning training, and improves the efficiency of model training and the reliability of the final safety performance assessment.
[0056] Furthermore, such as Figures 2-5As shown, using the method of this embodiment, the preset random seed is 20260304. For each vehicle type (VAN, SEDAN, MPV, SUV), 375 working condition samples are generated, including 160 core area samples, 185 global area samples (including 145 forced external samples (ring zone and outer area), and 40 soft transition samples), and 30 extreme value samples. Vehicle speed is set to a global range of [10,70] km / h and a core range of [30,50] km / h; road slope is set to a global range of [0,100] and a core range of [20,80]; collision angle is set to a global range of [-65,65]. The width of the side zone is set to 3.0 km / h in the vehicle speed direction and 6.0° in the slope direction. The ring zone width is 0.025 (normalized), the soft ring suppression strength is 2.2, and the minimum acceptance probability is 0.25. The optimized sampling candidate set size is 80, and the number of iterations is 220. The nearest neighbor distance threshold, coverage threshold, and repair rounds are all set according to actual engineering requirements.
[0057] Figure 2 In the diagram, (a) and (b) represent the distribution of vehicle speed parameters in the core area and the global area, respectively; (c) and (d) represent the distribution of road gradient parameters in the core area and the global area, respectively; and (e) and (f) represent the distribution of collision angle parameters in the core area and the global area, respectively. Figure 2 The two short dashed lines above (a) and (c) represent the lower and upper limits of the core area range of the corresponding parameters, respectively.
[0058] The following process is performed simultaneously and in parallel for the four vehicle models: First, generate core area samples according to step S120 above, ensuring that the boundaries are dense and without hollow areas; then, generate global samples in 8 sub-regions using the soft loop suppression probability model according to steps S130~S150, while supplementing soft transition samples to achieve smooth connection; then, generate 30 extreme value region samples with slight perturbations according to step S160; after merging, perform local repair and optimization; finally, perform multi-index verification, and if it does not meet the requirements, retry up to 120 times, retaining the result with the highest score.
[0059] Typical implementation results show that the minimum nearest neighbor distance across the entire domain is approximately 0.012, and the 95th percentile is approximately 0.168; the 95th percentile for external coverage is approximately 0.158; the proportion of ring-zone samples is approximately 0.085; the number of repair rounds is typically 0-2 rounds; and the success rate is 100%. Finally, an Excel file (filename doe_1500.xlsx) containing a total of 1500 working condition samples is generated. The table includes vehicle names such as "van," "sedan," "MPV," and "SUV," which can be directly imported into multi-rigid-body dynamics simulation software to quickly produce a high-quality machine learning training dataset.
[0060] When it is necessary to extend to higher dimensions (e.g., adding road surface adhesion coefficients or discrete vehicle characteristics), simply increase the normalization space to [0,1]. The corresponding adjustments to the partitioning and sampling strategy are sufficient; there is no need to change the core soft ring constraints and local repair mechanisms. All parameters in this implementation have been industrially validated in vehicle-pedestrian collision scenarios. Users can flexibly adjust parameters such as the total sample size N, core region proportion, dimension D, and ring zone width according to different training needs.
[0061] Example 2 like Figure 6 As shown, this embodiment provides a vehicle-pedestrian collision condition sample generation device, including: The sample quantity determination module 201 for each region is used to determine the sample quantity for each region based on the preset total sample quantity for the vehicle model. The sample quantity for each region includes core region samples, global region samples, and extreme value region samples. The core region includes an inner region and a side region. The global region includes a ring region and an outer region. The ring region is the region close to the core region, and the outer region is the region far away from the core region. The core area sample generation module 202 is used to generate the core area sample based on the number of core area samples and the volume ratio of each sub-region in the core area; The global region sample generation module 203 is used to determine the number of samples in each sub-region of the global region based on the number of global region samples and the volume ratio of each sub-region in the global region; calculate the acceptance probability of the ring candidate samples in the global region using a soft ring suppression probability model; determine the ring candidate samples as acceptable samples if the acceptance probability exceeds a preset probability; and generate the global region samples based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region. The extreme value region sample generation module 204 is used to generate the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations.
[0062] The device is used to perform the above method, and therefore has at least the functional modules and beneficial effects corresponding to the above method.
[0063] Example 3 like Figure 7 As shown, this embodiment provides an electronic device, including: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by at least one of the processors to enable the processor to perform the described method. Since at least one processor in the electronic device is capable of performing the described method, it thus possesses at least the same advantages as the described method.
[0064] Optionally, the electronic device also includes interfaces for connecting the various components, including high-speed interfaces and low-speed interfaces. The components are interconnected using different buses and can be mounted on a common motherboard or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI (Graphical User Interface) on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors can be used with multiple memories, and / or multiple buses can be used with multiple memories, if desired. Similarly, multiple electronic devices (e.g., as a server array, a group of blade servers, or a multiprocessor system) can be connected, each providing some of the necessary operations. Figure 7 Take processor 301 as an example.
[0065] The memory 302, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the vehicle-pedestrian collision condition sample generation method in this embodiment (e.g., the module for determining the number of samples in each region, the core region sample generation module, the global region sample generation module, and the extreme value region sample generation module in the vehicle-pedestrian collision condition sample generation device). The processor 301 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 302, thereby implementing the aforementioned vehicle-pedestrian collision condition sample generation method.
[0066] The memory 302 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on terminal usage. Furthermore, the memory 302 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 302 may further include memory remotely located relative to the processor 301, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0067] The electronic device may further include an input device 303 and an output device 304. The processor 301, memory 302, input device 303, and output device 304 can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0068] Input device 303 can receive input digital or character information, and output device 304 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor). The display device may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touchscreen.
[0069] Example 4 This embodiment provides a computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above. The computer instructions on this computer-readable storage medium, used to cause a computer to perform the methods described above, thus have at least the same advantages as the methods described above.
[0070] The medium in this application may be any combination of one or more computer-readable media. The medium may be a computer-readable signal medium or a computer-readable storage medium. The medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the medium (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, the medium may be any tangible medium containing or storing a program that may be used by or in connection with an instruction execution system, apparatus, or device.
[0071] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0072] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF (Radio Frequency), or any suitable combination thereof.
[0073] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0074] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0075] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for generating vehicle-pedestrian collision scenario samples, characterized in that, include: Based on the total preset sample size for the vehicle model, the number of samples in each region is determined. The samples in each region include core region samples, global region samples, and extreme value region samples. The core region includes an inner region and a side region. The global region includes a ring region and an outer region. The ring region is the region close to the core region, and the outer region is the region of the ring region away from the core region. The core area samples are generated based on the number of core area samples and the volume ratio of each sub-region in the core area; The number of samples in each sub-region of the global region is determined based on the number of samples in the global region and the volume ratio of each sub-region in the global region. The acceptance probability of ring candidate samples in the global region is calculated using a soft ring suppression probability model. If the acceptance probability exceeds a preset probability, the ring candidate sample is determined as an acceptable sample. The global region sample is generated based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region. The extreme value region sample is generated based on spatial corner points, surface center points, and high-risk combinations.
2. The method for generating vehicle-pedestrian collision condition samples according to claim 1, characterized in that, The process of generating the core region samples based on the number of core region samples and the volume ratio of each sub-region in the core region includes: Based on the number of samples in the core area and the volume ratio of each sub-region in the core area, the number of samples in the internal region and the number of samples in the side zone region are determined respectively. The core area samples are generated based on the number of samples in the internal region and the number of samples in the side zone region.
3. The method for generating vehicle-pedestrian collision condition samples according to claim 1, characterized in that, Before calculating the acceptance probability of ring candidate samples in the global region using the soft ring suppression probability model, the process also includes the step of constructing the soft ring suppression probability model, including: A soft ring suppression probability model is constructed based on the minimum acceptance probability, the normalized distance from the annular candidate sample to the core region, the annular region width, and the suppression intensity index.
4. The method for generating vehicle-pedestrian collision condition samples according to claim 3, characterized in that, The soft loop suppression probability model is as follows: ; Where p is the acceptance probability of the annular candidate sample, p_min is the minimum acceptance probability, dist is the normalized distance from the annular candidate sample to the core region, ringBand is the width of the annular region, and α is the suppression intensity index.
5. The method for generating vehicle-pedestrian collision condition samples according to claim 1, characterized in that, The step of generating the global region sample based on the sample quantity of each sub-region of the global region, the acceptable samples, and the parameter range of the core region includes: Based on the number of samples in each sub-region of the global region and the acceptable samples, a first global region sample is generated; Based on the parameter range of the core region, soft transition samples are generated, and all soft transition samples fall within the core region.
6. The method for generating vehicle-pedestrian collision condition samples according to claim 1, characterized in that, After generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes: Merge core region samples, global region samples, and extreme value region samples, and transform them into a normalized space; Based on the distance of each sample to the nearest other sample in the normalized space, the cluster point is determined. The cluster point is the sample point whose Euclidean distance to the nearest other sample point is less than a preset nearest neighbor distance threshold. For the cluster point, candidate samples are regenerated within the region to which the cluster point belongs, and the global external points are further screened and replaced using the soft loop suppression probability model until the distance constraint is met.
7. The method for generating vehicle-pedestrian collision condition samples according to claim 1, characterized in that, After generating the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations, the method further includes: Based on the generated samples from each region, the values of the evaluation indicators are calculated. The evaluation indicators include the minimum nearest neighbor distance for the entire region, the 95th percentile nearest neighbor distance, the 95th percentile of external spatial coverage, and the proportion of samples in the annular region. If the values of the evaluation indicators do not all meet the standards, a maximum number of retries is set, samples for each region are regenerated, and the result with the highest comprehensive score is selected as the final sample from the retries.
8. A device for generating vehicle-pedestrian collision scenario samples, characterized in that, include: The module for determining the number of samples in each region is used to determine the number of samples in each region based on the preset total number of samples for the vehicle model. The samples in each region include core region samples, global region samples, and extreme value region samples. The core region includes an inner region and a side region. The global region includes a ring region and an outer region. The ring region is the region close to the core region, and the outer region is the region far away from the core region. The core area sample generation module is used to generate the core area sample based on the number of core area samples and the volume ratio of each sub-region in the core area; The global region sample generation module is used to determine the number of samples in each sub-region of the global region based on the number of global region samples and the volume ratio of each sub-region in the global region. The acceptance probability of ring candidate samples in the global region is calculated using a soft ring suppression probability model. If the acceptance probability exceeds a preset probability, the ring candidate sample is determined as an acceptable sample. The global region sample is generated based on the number of samples in each sub-region of the global region, the acceptable samples, and the parameter range of the core region. The extreme value region sample generation module is used to generate the extreme value region sample based on spatial corner points, surface center points, and high-risk combinations.
9. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to at least one of the processors; The memory stores instructions executable by at least one of the processors, which are executed to enable the at least one of the processors to perform the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The medium stores computer instructions for causing the computer to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Sampling method suitable for advanced reactor continuous-discrete mixed variable design optimization
CN121031094A
Key area identification method based on four-dimensional representation index construction
CN122113449A