A sotif data collection method and related device
By screening and evaluating multi-source heterogeneous data at the vehicle end, the high cost of SOTIF data collection has been solved, achieving efficient data collection and screening and improving the data support capabilities of the intelligent driving system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GAC HONDA AUTOMOBILE CO LTD
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, SOTIF data acquisition requires a large amount of raw data backhaul and filtering work, resulting in high data acquisition costs and low efficiency, making it difficult to meet the data requirements of intelligent driving systems.
By calling multiple on-board environmental sensors to detect multi-frame, multi-source, heterogeneous data on the vehicle side, calculating the first Shannon entropy for filtering, and combining residual risk assessment, only data that meets the SOTIF conditions is uploaded.
It reduces reliance on closed venues and professional vehicle fleets, lowers data collection costs, improves the quality and efficiency of SOTIF data, and alleviates the processing burden on the cloud.
Smart Images

Figure CN122093767A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive technology, and in particular to a SOTIF data acquisition method and related equipment. Background Technology
[0002] With the development of automotive technology, the concept of functional safety and the ISO 26262 standard have become widely adopted in the automotive industry. ISO 26262 aims to address the safety risks arising from the failure of vehicle electronic and electrical systems. However, in the current booming development of the intelligent driving industry, intelligent driving systems that meet ISO 26262 still cannot provide consumers with sufficient peace of mind. For example, there have been numerous reported safety accidents in the market where vehicles had no electronic or electrical faults, yet still suffered accidents due to defects in their intelligent driving systems.
[0003] SOTIF can address potential hazards in vehicles caused by limitations or misuse of intended functions, thus offering higher safety than traditional automotive safety technologies. Because SOTIF identifies a broader range of automotive safety risks, it requires more data to implement (e.g., the training, testing, and validation data used to train an AI model capable of risk identification when supporting SOTIF), meaning SOTIF requires stronger data support.
[0004] Current technologies typically involve a professional fleet of vehicles driving in closed environments or other specific locations. High-risk test cases are pre-defined by personnel, and data from vehicle-mounted LiDAR point clouds and camera images are collected within these cases. All data is then transmitted back to the cloud, where it undergoes screening and labeling of high-risk driving scenarios in a laboratory environment. The final selected data is stored in a SOTIF database to support the construction of training datasets and the training of artificial intelligence models. However, because SOTIF requires a massive amount of data, and not all raw data collected by the fleet meets SOTIF requirements, a very large amount of data needs to be collected and transmitted back. The subsequent screening and labeling processes also present a significant workload, which is detrimental to the construction of the SOTIF database. Summary of the Invention
[0005] To address at least one of the aforementioned technical problems, the present invention aims to provide a SOTIF data acquisition method and related equipment.
[0006] On one hand, embodiments of the present invention include a SOTIF data acquisition method, the SOTIF data acquisition method comprising the following steps: The system utilizes multiple onboard environmental sensors to detect and obtain multiple frames of heterogeneous data from multiple sources. For any frame of the multi-source heterogeneous data, calculate the first Shannon entropy of the multi-source heterogeneous data; For any frame of the multi-source heterogeneous data, a judgment is made based on the first Shannon entropy of the multi-source heterogeneous data. When it is determined that the multi-source heterogeneous data belongs to a potential SOTIF scenario, the multi-source heterogeneous data is marked. Residual risk assessment is performed on the labeled multi-source heterogeneous data to obtain the residual risk value of the multi-source heterogeneous data; The multi-source heterogeneous data whose residual risk value is greater than the risk threshold are uploaded to the cloud.
[0007] Furthermore, the step of calling multiple on-board environmental sensors for detection to obtain multiple frames of multi-source heterogeneous data includes: Each of the vehicle-mounted environmental sensors is invoked to perform detection, and the original sensing data sequence detected by each of the vehicle-mounted environmental sensors is obtained; the original sensing data sequence includes multiple frames of original sensing data; The original sensor data sequences are synchronized in time so that the original sensor data of each frame are synchronized and aligned to the corresponding sampling time. For any frame of the original sensor data, perform initial target detection on the original sensor data to obtain traffic participant target information corresponding to the original sensor data; For any of the sampling times, target association matching and structured output are performed on all the traffic participant target information corresponding to the sampling time to obtain the multi-source heterogeneous data corresponding to the sampling time.
[0008] Further, calculating the first Shannon entropy of the multi-source heterogeneous data includes: Multiple detection dimensions are defined; the multiple detection dimensions include at least a semantic category dimension, a spatial grid dimension, and a velocity range dimension. Obtain the joint probability distribution of the multi-source heterogeneous data across all the detection dimensions; The first Shannon entropy is determined based on the joint probability distribution.
[0009] Further, obtaining the joint probability distribution of the multi-source heterogeneous data across all detection dimensions includes: Traverse all traffic participant targets in the multi-source heterogeneous data; For any of the traffic participant targets, determine the values of the traffic participant target in each detection dimension; Get target quantity ;in, This indicates that each of the aforementioned detection dimensions has a first... The first value, the first The first value... the first The number of traffic participant targets with each value; According to the formula
[0010]
[0011] The joint probability distribution is obtained by performing calculations. .
[0012] Further, determining the first Shannon entropy based on the joint probability distribution includes: According to the formula
[0013] Calculations are performed to obtain the first Shannon entropy. .
[0014] Further, the determination based on the first Shannon entropy of the multi-source heterogeneous data includes: Determine the baseline entropy; According to the formula
[0015] Perform calculations to determine the first entropy increment. ;in, This is the first Shannon entropy. The baseline entropy; Compare the first entropy increment with the entropy threshold. When the first entropy increment is greater than the entropy threshold, the multi-source heterogeneous data is determined to belong to a potential SOTIF scenario; otherwise, the multi-source heterogeneous data is determined not to belong to a potential SOTIF scenario.
[0016] Further, the determination based on the first Shannon entropy of the multi-source heterogeneous data includes: The system calls upon multiple onboard driving parameter sensors of the vehicle to detect and obtain multiple frames of driving parameter data. Determine a window time period; the end point of the window time period is the sampling time of the multi-source heterogeneous data; Obtain statistical feature quantities of the driving parameter data for multiple frames within the window time period; Obtain the second Shannon entropy; the second Shannon entropy is the Shannon entropy obtained with the starting point of the window time period as the sampling time. Obtain the second entropy increment of the first Shannon entropy relative to the second Shannon entropy; A judgment value is determined based on the second entropy increment and the statistical feature quantity; wherein the judgment value is positively correlated with the second entropy increment and negatively correlated with the statistical feature quantity. When the judgment value is greater than the judgment threshold, the multi-source heterogeneous data is determined to belong to a potential SOTIF scenario; otherwise, the multi-source heterogeneous data is determined not to belong to a potential SOTIF scenario.
[0017] Further, the residual risk assessment of the labeled multi-source heterogeneous data to obtain the residual risk value of the multi-source heterogeneous data includes: Define multiple sets of perturbation values; For any of the labeled multi-source heterogeneous data, perform multiple replay processes on the multi-source heterogeneous data; During any of the playback processes, a set of perturbation values is randomly selected to perturb the multi-source heterogeneous data. The perturbed multi-source heterogeneous data is then loaded into the digital twin environment for playback. The playback duration is recorded and monitored. When an unsafe event is detected, the number of unsafe events is accumulated. After all the playback processes are completed, the probability of unsafety in a single playback is determined based on the number of unsafe events and the total number of playback processes, and the average playback time of all playback processes is calculated. The quotient of the single playback insecurity probability and the average playback duration is calculated as the residual risk value.
[0018] On the other hand, embodiments of the present invention also include a computer device, including a memory and a processor, the memory for storing at least one program, and the processor for loading at least one program to execute the SOTIF data acquisition method in the embodiments.
[0019] On the other hand, embodiments of the present invention also include a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the SOTIF data acquisition method in the embodiments.
[0020] The beneficial effects of this invention are as follows: The SOTIF data acquisition method in the embodiments has a large data acquisition scope in the cloud, thereby reducing reliance on closed sites or professional data collection fleets and lowering the acquisition cost of multi-source heterogeneous data. Furthermore, First Automobile Works filters the multi-source heterogeneous data it collects using the first Shannon entropy method. The filtered multi-source heterogeneous data corresponds to traffic environments with high levels of disorder, meeting the basic application conditions of SOTIF. Moreover, First Automobile Works further performs residual risk assessment on the multi-source heterogeneous data that meets the first Shannon entropy condition, thereby further filtering. The filtered multi-source heterogeneous data corresponds to higher residual risk values, meeting the advanced application conditions of SOTIF. Therefore, by calling First Automobile Works to collect and upload multi-source heterogeneous data, it is easy to obtain a large amount of multi-source heterogeneous data that meets the application conditions of SOTIF, providing data support for SOTIF. Furthermore, the SOTIF data acquisition method is executed on the vehicle side, and by uploading the filtered multi-source heterogeneous data instead of the original full-volume data, the data processing burden on the cloud is significantly reduced, improving the implementation efficiency of SOTIF. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of a system where the SOTIF data acquisition method can be applied in the embodiment; Figure 2 This is a schematic diagram illustrating the steps of the SOTIF data acquisition method in the embodiment; Figure 3 This is a schematic diagram illustrating the principle of step S1 in the embodiment; Figure 4 This is a schematic diagram of the principle of step S302B in the embodiment. Detailed Implementation
[0022] Terminology Explanation: SOTIF: Safety Of The Intended Functionality, refers to technical standards that ensure intelligent vehicle systems (such as autonomous driving and advanced driver assistance systems) will not pose unreasonable risks due to limitations in system performance or foreseeable misuse under foreseeable and reasonable use conditions, within their intended functional scope. ISO 21448 defines relevant SOTIF standards. Compared to traditional automotive functional safety standards (such as ISO 26262), which focus on "system failure"—that is, random hardware failures or systematic software failures causing the system to "not function as designed"—SOTIF focuses on "performance limitations" and "functional deficiencies." While the system itself is not faulty and functions as designed, its design capabilities have limitations. Its core concept is "safety in unknown scenarios." GPU: Graphics Processing Unit; NPU: Neural Processing Unit.
[0023] This embodiment provides a SOTIF data acquisition method. The SOTIF data acquisition method can be applied to… Figure 1 The system shown is a system installed on a specific car, namely the first car. (Refer to...) Figure 1 The system includes a control module, a communication module, multiple on-board environmental sensors, and multiple on-board driving parameter sensors. The control module is a component with data acquisition, processing, and control functions; each on-board environmental sensor is used to detect environmental parameters, for example... Figure 1 The system includes three onboard environmental sensors: a visible light camera, a lidar, and a millimeter-wave radar. These sensors detect environmental parameters such as visible light images, lidar point clouds, and millimeter-wave point clouds. Various onboard driving parameter sensors are used to detect driving parameters, for example... Figure 1 The system includes three onboard driving parameter sensors: a vehicle speed sensor, a satellite navigation module, and a steering wheel angle sensor. These sensors detect driving parameters such as vehicle speed, real-time vehicle position, and steering wheel angle. (Refer to...) Figure 1 The communication module can communicate wirelessly with the cloud server.
[0024] In this embodiment, the control module at the vehicle end of the first vehicle can execute the various steps of the SOTIF data acquisition method.
[0025] Reference Figure 2 The SOTIF data acquisition method includes the following steps: S1. Call multiple on-board environmental sensors of this vehicle to detect and obtain multiple frames of multi-source heterogeneous data; S2. For any frame of multi-source heterogeneous data, calculate the first Shannon entropy of the multi-source heterogeneous data; S3. For any frame of multi-source heterogeneous data, make a judgment based on the first Shannon entropy of the multi-source heterogeneous data. When it is determined that the multi-source heterogeneous data belongs to a potential SOTIF scene, mark the multi-source heterogeneous data. S4. Perform residual risk assessment on the labeled multi-source heterogeneous data to obtain the residual risk value of the multi-source heterogeneous data; S5. Upload the multi-source heterogeneous data whose corresponding residual risk value is greater than the risk threshold to the cloud.
[0026] The principle of step S1 is as follows: Figure 3 As shown. (Refer to...) Figure 3The control module of the first automobile can set a sampling time at regular intervals (e.g., 100ms) to set... , , ...and so on, multiple sampling times. For the visible light camera, a vehicle-mounted environmental sensor, its sampling times... A single frame of raw sensor data is captured, which is visible light image 1, at the sampling time. A single frame of raw sensor data, which is known as visible light image 2, is captured at the sampling time. A single frame of raw sensor data is captured, resulting in a visible light image 3... thus forming a raw sensor data sequence 1 such as visible light image 1, visible light image 2, visible light image 3... Similarly, for a vehicle-mounted environmental sensor like LiDAR, it detects a raw sensor data sequence 2 such as laser point cloud 1, laser point cloud 2, laser point cloud 3... For a vehicle-mounted environmental sensor like millimeter-wave radar, it detects a raw sensor data sequence 3 such as millimeter-wave point cloud 1, millimeter-wave point cloud 2, millimeter-wave point cloud 3...
[0027] After obtaining the raw sensor data sequences, they can be preprocessed. For example, for each visible light image, image denoising and distortion correction can be performed; for each laser point cloud, isolated points and ground points can be removed; and for each millimeter-wave point cloud, Kalman filtering and velocity jitter elimination can be performed.
[0028] Reference Figure 3 The original sensor data sequences, including original sensor data sequence 1, original sensor data sequence 2, and original sensor data sequence 3, are synchronized in time to ensure that each frame of original sensor data is aligned to its corresponding sampling time. For example, visible light image 1, laser point cloud 1, and millimeter-wave point cloud 1 are all aligned to their respective sampling times. .
[0029] In this embodiment, spatial registration can also be performed on each of the original sensor data sequences, such as original sensor data sequence 1, original sensor data sequence 2, and original sensor data sequence 3. For example, based on the calibration parameters (intrinsic and extrinsic parameters) of the installation locations of the visible light camera, LiDAR, and millimeter-wave radar, the LiDAR point cloud and millimeter-wave point cloud are uniformly mapped to the image coordinate system or vehicle coordinate system where the visible light image is located. In this way, pixels at the same location in the environment within the visible light image, LiDAR point cloud, and millimeter-wave point cloud can be mutually mapped.
[0030] Reference Figure 3In step S1, after obtaining various types of raw sensor data collected at each sampling time, preliminary target detection can be performed on the raw sensor data to obtain traffic participant target information corresponding to the raw sensor data. For example, for the data collected at each sampling time... The detected visible light image 1, laser point cloud 1, and millimeter-wave point cloud 1 can be analyzed using the YOLO / Transformer model. This allows for the detection of traffic participants such as pedestrians, vehicles, and cyclists in the visible light image 1, including their coordinates in the image plane. Similarly, the laser point cloud 1 can be analyzed using point cloud clustering and classification models, revealing the three-dimensional spatial coordinates of pedestrians, vehicles, and cyclists within it. Finally, the millimeter-wave point cloud 1 can be analyzed using point cloud clustering and classification models, detecting the radial distance and radial velocity of pedestrians, vehicles, and cyclists relative to the vehicle itself.
[0031] After obtaining the traffic participant target information corresponding to each raw sensor data point, target association matching can be performed on all traffic participant target information corresponding to the same sampling time. For example, for sampling time... The detected visible light image 1, laser point cloud 1, and millimeter-wave point cloud 1 are used to identify the same traffic participant targets contained in these three images. The three-dimensional spatial coordinates, radial distance, and radial velocity corresponding to the same traffic participant targets are then organized and mapped to these targets, thus forming a structured output. For example, in this embodiment, the traffic participant target information of one of the traffic participant targets detected at a certain sampling time can be output in the form of attribute fields as shown in Table 1.
[0032] Table 1
[0033] Based on sampling time For example, sampling time Each detected traffic participant target information can be converted into the attribute fields shown in Table 1. If there are multiple traffic participant target information, there will be multiple attribute fields as shown in Table 1, thus forming the sampling time. The corresponding multi-source heterogeneous data. For , ...The same processing can be performed at other sampling times to obtain multi-source heterogeneous data corresponding to other sampling times.
[0034] In this embodiment, when the control module executes step S2, which is to calculate the first Shannon entropy of the multi-source heterogeneous data, it can specifically perform the following steps: S201. Determine multiple detection dimensions; S202. Obtain the joint probability distribution of multi-source heterogeneous data across all detection dimensions; S203. Determine the first Shannon entropy based on the joint probability distribution.
[0035] Step S2, namely steps S201-S203, is performed for each frame of multi-source heterogeneous data. In this embodiment, it is based on the sampling time. The following explanation uses the execution steps S201-S203 of a corresponding frame of multi-source heterogeneous data as an example.
[0036] In this embodiment, step S201 can set three detection dimensions: semantic category dimension, spatial grid dimension, and velocity range dimension. More dimensions can also be set.
[0037] Specifically, the semantic category dimension represents the set of categories to which all traffic participant target information belongs in multi-source heterogeneous data. For example, the semantic category dimension can be represented as { (pedestrian), (Small car) (Large trucks) (Cyclist) (Obstacles)}; Spatial grid dimension represents the grid division of the spatial location of all traffic participants' target information in multi-source heterogeneous data. For example, the grid can be divided according to the vehicle's body coordinate system. For instance, a 100m × 30m area in front of the vehicle can be divided into 20 × 6 = 120 grids of 5m × 5m size. Thus, the spatial grid dimension can be represented as { (First grid) (Grid 1) (3rd grid)...}; The speed interval dimension represents the set of target information of all traffic participants in multi-source heterogeneous data relative to the speed (or speed interval) of the vehicle. For example, the speed interval dimension can be represented as { (0-5m / s) (5-15m / s) (15-25m / s) (>25m / s)}.
[0038] In step S202, for the sampling time For a given frame of multi-source heterogeneous data, traverse all traffic participant targets within it and determine the value of each traffic participant target in the semantic category dimension. Values in the spatial grid dimension and the value in the velocity range dimension For example, if the sampling time In a corresponding frame of multi-source heterogeneous data, there is a pedestrian located in the nearest grid cell to the left of the vehicle, with a speed of 13 m / s relative to the vehicle. What is the semantic category value of this traffic participant target? (Pedestrian), the value in the spatial grid dimension is (First grid), the value in the velocity interval dimension is That is, for this one traffic participant objective, =1、 =1、 =1; for other traffic participants' objectives, there may be other corresponding objectives. , and The possible combinations of values.
[0039] In step S202, for the sampling time The corresponding frame of multi-source heterogeneous data can be used to statistically determine the number of detection dimensions. The first value, the first The first value... the first The number of traffic participant targets with each value Since this embodiment specifically sets three detection dimensions, step S202 can detect that each detection dimension has a first value. The first value, the first The first value, the first Each value (i.e., the value in the semantic category dimension) is... The value in the spatial grid dimension is The value in the velocity range dimension is The number of traffic participants is .
[0040] In step S202, traverse all (Specifically in this embodiment) , , ),right To perform summation, that is, according to the formula (1) Perform calculations to obtain the total number of states. Total number of states The meaning is sampling time. The size of the state space occupied by all traffic participants in each detection dimension in a corresponding frame of multi-source heterogeneous data.
[0042] In step S202, for any (Specifically in this embodiment) , , ), can be based on the formula (2) (Specifically in this embodiment) ) Perform calculations to obtain the joint probability distribution. (Specifically in this embodiment) ).
[0044] In step S203, traverse all (Specifically in this embodiment) , , According to the formula (3) Calculations were performed to obtain the first Shannon entropy. .in, Logarithmic operations with base 2 are mathematically considered as... =0 It is meaningless, but in this embodiment, it is defined as when =0 =0, thus enabling the calculation of the first Shannon entropy. .
[0046] In this embodiment, the first Shannon entropy The meaning is a specific sampling time (e.g., the sampling time selected in this embodiment). In a frame of multi-source heterogeneous data collected, the entropy value of the state of all traffic participants represents the degree of order of all traffic participants in the traffic environment where the first car is located at a specific sampling time. Specifically, the first Shannon entropy... The larger the scale, the lower the level of order, and the higher the degree of confusion among all traffic participants. According to the objective laws of road traffic, the greater the traffic risks faced by the first car, such as traffic accidents.
[0047] In this embodiment, based on the execution of steps S201-S203, when the control module executes step S3, which is the step of making a judgment based on the first Shannon entropy of the multi-source heterogeneous data, the following steps can be specifically executed: S301A. Determine the baseline entropy; S302A. According to the formula (4) Perform calculations to determine the first entropy increment. ; S303A. Compare the magnitude of the first entropy increment with the entropy threshold; S304A. When the first entropy increment is greater than the entropy threshold, the multi-source heterogeneous data is determined to belong to the potential SOTIF scenario; otherwise, the multi-source heterogeneous data is determined not to belong to the potential SOTIF scenario.
[0049] Steps S301A-S304A are the first execution method of step S3.
[0050] In step S301A, a fixed entropy value can be determined as the baseline entropy. In this embodiment, baseline entropy For a car in a normal, orderly traffic environment (e.g., an open road, a constant-speed traffic flow), through the first Shannon entropy... The entropy value calculated using the same principle [e.g., formulas (1)-(3)] can be obtained through calibration experiments, etc. In this embodiment, the baseline entropy... Its size is 0.8 bits.
[0051] In step S302A, the first Shannon entropy obtained from steps S201-S203 is used. and the baseline entropy obtained by performing step S301A The first entropy increment is calculated according to formula (4). First entropy increment The meaning is that the first vehicle is at a specific sampling time (e.g., the sampling time selected in this embodiment). The degree of chaos in the actual traffic environment for all traffic participants, and the deviation from the degree of chaos in a normal and orderly traffic environment.
[0052] In step S303A, a fixed-size entropy value (e.g., 0.35 bits) can be set as the entropy threshold. The first entropy increment is then... Compare the magnitude with the entropy threshold. If the first entropy increment... If the increment is greater than the entropy threshold, then the first entropy increment is considered to be... Larger, meaning the first vehicle at a specific sampling time (e.g., the sampling time selected in this embodiment). The actual traffic environment exhibits a high degree of confusion among all traffic participants, thus affecting the accuracy of this sampling time (e.g., sampling time). The collected multi-source heterogeneous data is determined to belong to a potential SOTIF scenario; otherwise, the first entropy increment is considered to be... Smaller, relatively lower level of confusion among all traffic participants' goals, this specific sampling time (e.g., sampling time) The collected multi-source heterogeneous data was determined not to belong to the potential SOTIF scenario.
[0053] In this embodiment, based on the execution of steps S201-S203, when the control module executes step S3, which is the step of making a judgment based on the first Shannon entropy of the multi-source heterogeneous data, the following steps can be specifically executed: S301B. Calls multiple on-board driving parameter sensors of this vehicle to detect and obtain multiple frames of driving parameter data; S302B. Determine the window time period; S303B. Obtain statistical features of driving parameter data for multiple frames within a window time period; S304B. Obtain the second Shannon entropy; S305B. Obtain the second entropy increment of the first Shannon entropy relative to the second Shannon entropy; S306B. Determine the judgment value based on the second entropy increment and statistical characteristic quantities; S307B. When the judgment value is greater than the judgment threshold, the multi-source heterogeneous data is judged to belong to the potential SOTIF scenario; otherwise, the multi-source heterogeneous data is judged not to belong to the potential SOTIF scenario.
[0054] Steps S301B-S307B are the second execution method of step S3. That is, when executing step S3, you can choose not to execute steps S301A-S304A and instead choose to execute steps S301B-S307B.
[0055] In step S301B, similar to the principle of calling various on-board environmental sensors, the control module can... , , ...At each sampling time, multiple onboard driving parameter sensors, such as the vehicle speed sensor, satellite navigation module, and steering wheel angle sensor, are used to detect and obtain multiple frames of driving parameter data. For example, at each sampling time... The vehicle speed sensor detected the vehicle's speed as being [missing information]. The satellite navigation module detected the vehicle's real-time location (latitude and longitude coordinates) as ( , The steering wheel angle sensor detects the steering wheel's rotation angle (0 for maintaining straight-line driving, negative for left turns, and positive for right turns). Thus obtaining the sampling time Corresponding driving parameter data for one frame =( , , , ).
[0056] In step S302B, refer to Figure 4 The sampling time corresponding to the multi-source heterogeneous data targeted in the current execution step S3 (in this embodiment, the sampling time) Set the duration as the endpoint. (Specifically, it can be a fixed value, such as 10s), that is, from the sampling time. Start Prep Duration This determines the window time period. .
[0057] Window period The starting point is a specific sampling time (in this embodiment, the sampling time). ), window time period At least the sampling time is included. and sampling time These two sampling times can also include more sampling times.
[0058] In step S303B, for the window time period At each sampling time point within the timeframe, acquire the driving parameter data for each frame collected at these sampling times, including at least the sampling time points. A frame of driving parameter data collected =( , , , For these multi-frame driving parameter data, the window time period can be calculated. The corresponding first frame of driving parameter data (i.e., in this embodiment) ) and the last frame of driving parameter data (i.e., in this embodiment) The magnitude of the vector difference between the vectors is used as the statistical characteristic quantity to be obtained in step S303B. It can also calculate window time periods. The difference between the maximum and minimum values of the vector moduli of all corresponding driving parameter data in each frame is used as the statistical feature quantity to be obtained in step S303B. It can also calculate window time periods. The variance (or standard deviation) of the vector magnitude of all corresponding driving parameter data for each frame is used as the statistical feature quantity to be obtained in step S303B. .
[0059] Various forms of statistical characteristic quantities are obtained by performing step S303B. All of these can represent the window time period. Within, the degree of drastic change in the driving parameters of the first car, specifically, statistical characteristics. The larger the value, the higher the driving parameters of the first car within the window time period. The more drastic the internal changes.
[0060] In step S304B, the starting point of the window time period, i.e., the sampling time, is obtained. The corresponding Shannon entropy is used as the second Shannon entropy. Specifically, at the sampling time At that time, the control module can calculate the sampling time according to the principle of steps S201-S203. The corresponding "first Shannon entropy" is stored locally, thus serving as the second Shannon entropy in step S304B. Second Shannon entropy The meaning is the starting point of the first vehicle's time window, i.e., the sampling time. The degree of confusion of the goals of all traffic participants in the actual traffic environment.
[0061] In step S305B, the formula can be used. (5) Perform calculations to determine the second entropy increment. .
[0063] In step S306B, the second entropy increment calculated based on step S305B is... The statistical characteristic quantity calculated in step S303B Calculate the second entropy increment Positive correlation, statistical characteristics negative correlation judgment value Specifically, in this embodiment, it can be based on the formula (6) Perform calculations to determine the judgment value. .
[0065] In step S307B, a fixed-size judgment threshold (e.g., 2) can be set. The judgment value... Compare the value with the judgment threshold. If the judgment value... If the value is greater than the judgment threshold, then the judgment value is considered to be... Larger, thus reducing this specific sampling time (e.g., sampling time) The collected multi-source heterogeneous data is determined to belong to a potential SOTIF scenario; otherwise, the judgment value is considered to be... Smaller, this specific sampling time (e.g., sampling time) The collected multi-source heterogeneous data was determined not to belong to the potential SOTIF scenario.
[0066] In this embodiment, the principle of executing steps S301B-S307B is as follows: By executing steps S301B-S307B, not only is the degree of confusion of the goals of all traffic participants in the traffic environment where the first car is located considered, but also the impact of the first car itself as a traffic participant on the degree of confusion of the goals of traffic participants. For example, when the rate of change of the driving parameters of the first car is small, the driving of the first car is relatively smooth, and the first car itself has little impact on the traffic environment. Conversely, when the rate of change of the driving parameters of the first car is large, the driving of the first car is relatively aggressive, and the first car itself has a greater impact on the traffic environment. Therefore, the statistical feature quantity obtained by executing steps S301B-S303B is... This represents the rate of change of the driving parameters of the first vehicle, while the second entropy increment obtained by executing steps S304B-S305B is... This indicates the change in the degree of confusion of the goals of all traffic participants in the traffic environment where the first car is located within the same window time period; the second entropy increment is calculated by executing step S306B. Positively correlated with statistical characteristic negative correlation judgment value It can obtain the second entropy increment. Judgment value for maintaining the same trend Furthermore, this offsets at least part of the impact of the first vehicle's driving changes on the confusion of all traffic participants' goals in the traffic environment where the first vehicle is located. This means that when performing step S307B, the judgment of whether multi-source heterogeneous data belongs to a potential SOTIF scenario is less affected by the first vehicle's own driving behavior. In other words, the judgment result is closer to the information contained in the multi-source heterogeneous data itself, allowing the first vehicle to exist primarily as an observer. Thus, the multi-source heterogeneous data uploaded to the cloud has lower specificity for the first vehicle, resulting in better generality and providing more comprehensive data support for the implementation of SOTIF. Moreover, during the execution of S301B-S307B, it is only necessary to calculate the "first Shannon entropy" detected by the first vehicle in the traffic environment at each sampling time, without using the baseline entropy in steps S301A-S304A. Such a specific fixed value, thereby reducing the baseline entropy Inaccurate detection can lead to biases in the determination of attributes of multi-source heterogeneous data.
[0067] In this embodiment, since the multi-source heterogeneous data marked in step S3 belongs to the potential SOTIF scenario, it is sufficient to provide data support for the implementation of SOTIF. Therefore, after completing steps S1-S3, step S4 can be skipped, and the multi-source heterogeneous data marked in step S3, that is, the multi-source heterogeneous data belonging to the potential SOTIF scenario, can be directly uploaded to the cloud.
[0068] In this embodiment, if the computing power of the first vehicle is sufficient, step S4 can be further executed based on steps S1-S3. Specifically, when the control module executes step S4, which is to perform residual risk assessment on the marked multi-source heterogeneous data and obtain the residual risk value of the multi-source heterogeneous data, the following steps can be performed: S401. Define multiple sets of disturbance values; S402. For any labeled multi-source heterogeneous data, perform multiple replay processes on the multi-source heterogeneous data; S403. During any playback process, a set of perturbation values is randomly selected to perturb the multi-source heterogeneous data. The perturbed multi-source heterogeneous data is then loaded into the digital twin environment for playback. The playback duration is recorded and monitored. When an unsafe event is detected, the number of unsafe events is accumulated. S404. After all playback processes are completed, determine the probability of unsafety in a single playback based on the number of unsafe events and the total number of playback processes, and calculate the average playback time of all playback processes; S405. Calculate the quotient of the unsafe probability of a single playback and the average playback duration as the residual risk value.
[0069] In step S401, the control module defines multiple sets of disturbance values. For example, each set of disturbance values is a combination of specific values such as "the proportion of missing lidar point clouds, the road adhesion coefficient, and the deviation of acceleration and deceleration of adjacent vehicles".
[0070] In step S402, for any multi-source heterogeneous data marked in step S3, multiple replay processes are performed on the multi-source heterogeneous data. Specifically, the control module can call upon the on-board GPU computing power of the first vehicle to perform a large number of (specifically...) =10 6 The replay process of (number) times.
[0071] During any playback, a set of disturbance values is randomly selected (e.g., "LiDAR point cloud missing percentage is 5%, road adhesion coefficient is +0.5, and adjacent vehicle acceleration / deceleration deviation is +1.5 m / s"). 2The multi-source heterogeneous data selected in step S402 is perturbed, and the perturbed multi-source heterogeneous data is loaded into the digital twin environment for playback. Specifically, during playback in the digital twin environment, the duration of the playback process is recorded, and simulation is performed based on the perturbed multi-source heterogeneous data. It is determined whether any unsafe events are found during the playback process, such as "collision risk [e.g., the minimum distance between the vehicle and the target is less than the safety threshold (e.g., 3m for a car, 1.5m for a pedestrian)]", "system failure [e.g., missed / false detection of perception (e.g., failure to identify a pedestrian), decision error (e.g., misjudging that another vehicle is not changing lanes), or planned trajectory crossing the line]", or "emergency operation [e.g., triggering safety systems such as AEB (Automatic Emergency Braking) or ESC (Electronic Stability Control)]". If one or more unsafe events occur, this playback process is marked as "1", and the number of unsafe events is calculated according to the actual number of unsafe events detected. Accumulate the values; otherwise mark them as "0".
[0072] After the entire playback process is completed, step S404 is executed to calculate the final accumulated number of unsafe events. The total number of times the playback process was repeated ( =10 6 Determine the probability of insecurity in a single playback. For example, it can be based on the formula (7) Calculations are performed to obtain the probability of insecurity for a single playback. .
[0074] In step S404, the control module also calculates the average duration of each playback process, i.e., the average playback duration. .
[0075] In step S405, the probability of insecurity for a single playback is calculated. With average playback duration The quotient is used as the residual risk value, i.e., according to the formula (8) The residual risk value was calculated. In this embodiment, the residual risk value The meaning is the probability of an unsafe event occurring within a unit of time, determined by replaying a digital twin model based on labeled multi-source heterogeneous data.
[0077] In this embodiment, a fixed risk threshold can be set (e.g., 10). -5 h -1In step S5, the residual risk value corresponding to each multi-source heterogeneous data is determined. The relationship between risk and risk threshold. If the residual risk value corresponding to multi-source heterogeneous data... If the risk exceeds the risk threshold, the multi-source heterogeneous data is determined to have high residual risk and is uploaded to the cloud; otherwise, it is determined to have low residual risk and is discarded instead of being uploaded to the cloud. In this way, the multi-source heterogeneous data received by the cloud will all belong to potential SOTIF scenarios, or be both potential SOTIF scenarios and high residual risk multi-source heterogeneous data.
[0078] In this embodiment, when the control module executes step S5, it can encrypt the multi-source heterogeneous data that needs to be uploaded, and send the encrypted multi-source heterogeneous data to the communication module, which then sends the multi-source heterogeneous data to the cloud server.
[0079] In this embodiment, the first car executing steps S1-S5 is any car, that is, each car can execute steps S1-S5 to upload multi-source heterogeneous data that has been collected, processed and filtered by the vehicle end to the cloud, so that the cloud can receive multi-source heterogeneous data from different cars.
[0080] Furthermore, the first vehicle executing steps S1-S5 can be a dedicated data sampling vehicle, or a commercial vehicle, a private car, or other vehicle not specifically designed for data sampling. These vehicles can execute steps S1-S5 during actual driving or operation. Therefore, the cloud can obtain multi-source heterogeneous data uploaded by various types of vehicles, thereby expanding the scope of data collection.
[0081] A SOTIF database is established on the cloud server to store the multi-source heterogeneous data uploaded by First Automobile Works. The cloud server can also be configured with knowledge graph update units and federated learning coordinators to process the received multi-source heterogeneous data and provide data support for the implementation of SOTIF.
[0082] In this embodiment, by executing steps S1-S5, the cloud has a larger data collection surface, thereby reducing reliance on closed sites or professional data collection fleets and lowering the cost of collecting multi-source heterogeneous data. Furthermore, First Automobile Works filters the collected multi-source heterogeneous data using the first Shannon entropy method. The filtered multi-source heterogeneous data corresponds to traffic environments with high levels of disorder, meeting the basic application conditions of SOTIF. Moreover, First Automobile Works further conducts residual risk assessments on the multi-source heterogeneous data that meets the first Shannon entropy conditions, thereby further filtering. The filtered multi-source heterogeneous data corresponds to... The relatively high residual risk value meets the advanced application conditions of SOTIF. Therefore, by calling First Automobile Works to collect and upload multi-source heterogeneous data, it is easy to obtain a large amount of multi-source heterogeneous data that meets the application conditions of SOTIF, providing data support for SOTIF. Moreover, steps S1-S5 are executed on the vehicle side, and by uploading the filtered multi-source heterogeneous data instead of the original full data, the local computing power of the vehicle's GPU, NPU and other local computing power can be fully utilized, reducing the waste caused by the idle local computing power of the vehicle side, greatly reducing the data processing burden on the cloud, and improving the implementation efficiency of SOTIF.
[0083] A computer program that executes the SOTIF data acquisition method in this embodiment can be written into a computer device or storage medium. When the computer program is read out and run, the SOTIF data acquisition method and / or the SOTIF data acquisition method in this embodiment can be executed, thereby achieving the same technical effect as the SOTIF data acquisition method and / or the SOTIF data acquisition method in the embodiment.
[0084] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. Furthermore, the descriptions of "upper," "lower," "left," and "right" used in this disclosure are only relative to the relative positional relationships of the components of this disclosure in the accompanying drawings. The singular forms "a," "an," and "the" used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. Moreover, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this embodiment specification is only for describing particular embodiments and is not intended to limit the invention. The term "and / or" as used in this embodiment includes any combination of one or more of the associated listed items.
[0085] It should be understood that although the terms first, second, third, etc., may be used to describe various elements in this disclosure, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, a first element may also be referred to as a second element without departing from the scope of this disclosure, and similarly, a second element may also be referred to as a first element. The use of any and all instances or exemplary language (“e.g.,” “such as,” etc.) provided in this embodiment is intended only to better illustrate embodiments of the invention and, unless otherwise required, does not impose a limitation on the scope of the invention.
[0086] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium. The method can be implemented using standard programming techniques—including a non-transitory computer-readable storage medium configured with a computer program, wherein such a storage medium causes the computer to operate in a specific and predefined manner—according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed application-specific integrated circuit (ASIC).
[0087] Furthermore, the procedures described in this embodiment can be performed in any suitable order unless otherwise indicated by this embodiment or otherwise obviously contradict the context. The procedures (or variations and / or combinations thereof) described in this embodiment can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. A computer program includes a plurality of instructions executable by one or more processors.
[0088] Furthermore, the method can be implemented in any suitable type of computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices, etc. Aspects of the invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. Furthermore, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. The invention of this embodiment includes these and other different types of non-transitory computer-readable storage media when such media comprises instructions or programs that implement the steps above in conjunction with a microprocessor or other data processor. When programmed according to the methods and techniques of the invention, the invention also includes the computer itself.
[0089] A computer program can be applied to input data to perform the functions of this embodiment, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices, such as a display. In a preferred embodiment of the invention, the transformed data represents physical and tangible objects, including specific visual depictions of physical and tangible objects generated on the display.
[0090] The above are merely preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention, as long as they achieve the technical effects of the present invention by the same means, should be included within the scope of protection of the present invention. Within the scope of protection of the present invention, the technical solutions and / or implementation methods can have various modifications and variations.
Claims
1. A SOTIF data acquisition method, executed at the vehicle end of a first automobile, characterized in that, The SOTIF data acquisition method comprises: calling a plurality of vehicle-mounted environment sensors of the vehicle to detect and obtain a plurality of frames of multi-source heterogeneous data; for any frame of the multi-source heterogeneous data, calculating a first Shannon entropy of the multi-source heterogeneous data; for any frame of the multi-source heterogeneous data, judging according to the first Shannon entropy of the multi-source heterogeneous data, and when it is determined that the multi-source heterogeneous data belongs to a potential SOTIF scene, marking the multi-source heterogeneous data; performing residual risk assessment on the marked multi-source heterogeneous data to obtain a residual risk value of the multi-source heterogeneous data; uploading the multi-source heterogeneous data corresponding to the residual risk value greater than a risk threshold to the cloud.
2. The SOTIF data collection method of claim 1, wherein, The calling of the plurality of vehicle-mounted environment sensors to detect and obtain a plurality of frames of multi-source heterogeneous data comprises: respectively calling each of the vehicle-mounted environment sensors to detect and obtain a raw sensor data sequence detected by each of the vehicle-mounted environment sensors; the raw sensor data sequence comprises a plurality of frames of raw sensor data; time synchronizing each of the raw sensor data sequences to synchronize and align each frame of the raw sensor data to a corresponding sampling time point; for any frame of the raw sensor data, performing target preliminary detection on the raw sensor data to obtain target information of a traffic participant corresponding to the raw sensor data; for any sampling time point, performing target association matching and structured output on all the target information of the traffic participant corresponding to the sampling time point to obtain the multi-source heterogeneous data corresponding to the sampling time point.
3. The SOTIF data collection method of claim 2, wherein, The calculation of the first Shannon entropy of the multi-source heterogeneous data comprises: determining a plurality of detection dimensions; the plurality of detection dimensions at least comprise a semantic category dimension, a spatial grid dimension and a speed interval dimension; obtaining a joint probability distribution of the multi-source heterogeneous data in all the detection dimensions; determining the first Shannon entropy according to the joint probability distribution.
4. The SOTIF data collection method of claim 3, wherein, The obtaining of the joint probability distribution of the multi-source heterogeneous data in all the detection dimensions comprises: traversing all the target information of the traffic participant in the multi-source heterogeneous data; for any target information of the traffic participant, determining a value of the target information of the traffic participant in each detection dimension; Acquisition target quantity ; wherein, represents the number of the traffic participant targets having the first value, the first value, …, the first value in each of the detection dimensions, respectively; according to a formula performing a calculation to obtain the joint probability distribution .
5. The SOTIF data collection method of claim 4, wherein, The determination of the first Shannon entropy according to the joint probability distribution comprises: according to a formula performing a calculation to obtain the first Shannon entropy .
6. The SOTIF data collection method of any one of claims 3-5, wherein, The judgment according to the first Shannon entropy of the multi-source heterogeneous data comprises: determining a baseline entropy; according to a formula calculations are performed to determine a first entropy increment ; wherein, is the first Shannon entropy, is the baseline entropy; comparing the first entropy increment with an entropy threshold in size; when the first entropy increment is greater than the entropy threshold, determining that the multi-source heterogeneous data belongs to a potential SOTIF scene, otherwise, determining that the multi-source heterogeneous data does not belong to a potential SOTIF scene.
7. The SOTIF data collection method of any one of claims 3-5, wherein, The judgment according to the first Shannon entropy of the multi-source heterogeneous data comprises: calling a plurality of vehicle-mounted driving parameter sensors of the vehicle to detect and obtain a plurality of frames of driving parameter data; determining a window time period; an end point of the window time period is a sampling time point of the multi-source heterogeneous data; obtaining statistical characteristic quantities of the corresponding plurality of frames of the driving parameter data in the window time period; obtaining a second Shannon entropy; the second Shannon entropy is a Shannon entropy obtained by taking the start of the window time period as a sampling time; obtaining a second entropy increment of the first Shannon entropy relative to the second Shannon entropy; determining a judgment value according to the second entropy increment and the statistical characteristic quantity; wherein the judgment value is positively correlated with the second entropy increment and negatively correlated with the statistical characteristic quantity; when the judgment value is greater than the judgment threshold, determining that the multi-source heterogeneous data belongs to a potential SOTIF scenario, otherwise, determining that the multi-source heterogeneous data does not belong to a potential SOTIF scenario.
8. The SOTIF data collection method of claim 3, wherein, the residual risk assessment of the labeled multi-source heterogeneous data to obtain a residual risk value of the multi-source heterogeneous data, comprising: defining a plurality of perturbation values; for any labeled multi-source heterogeneous data, performing a plurality of playback processes on the multi-source heterogeneous data; in any of the playback processes, randomly selecting a set of perturbation values to perturb the multi-source heterogeneous data, loading the perturbed multi-source heterogeneous data into a digital twin environment for playback, recording and monitoring the playback duration of the playback process, and when an unsafe event is monitored, accumulating the number of unsafe event occurrences; after all the playback processes are completed, determining a single playback unsafe probability according to the number of unsafe event occurrences and the total number of playback processes, and calculating the average playback duration of all playback processes; calculating the quotient of the single playback unsafe probability and the average playback duration as the residual risk value.
9. A computer apparatus, comprising: comprising a memory and a processor, the memory being used to store at least one program, and the processor being used to load the at least one program to execute the SOTIF data acquisition method of any one of claims 1-8.
10. A computer readable storage medium having stored therein a program which is executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute the SOTIF data acquisition method of any one of claims 1-8.