Camera obstacle screening method and system for intelligent driving car
Patent Information
- Application Number
- CN202611134882.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明提供一种智能驾驶汽车用摄像头障碍物筛选方法,其主要目的在于解决障碍物检测准确率不高的问题
[0008]本发明实施例中,通过置信度驱动的双路径差异化筛选机制,实现了对潜在障碍物的精准过滤与可信度增强。具体而言,通过高低置信度的分流处理,避免了对所有目标统一执行高开销的运动跟踪与威胁评估,低置信度集合的二次验证机制有效保留了远距离小目标、部分遮挡目标等困难样本,提升了系统在复杂环境下的检测召回率,同时实现了一种能够针对不同置信度的潜在障碍物进行分层差异化筛选,进而提高了障碍物检测准确率。
Smart Images

Figure CN122842084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to a method and system for obstacle screening using a camera in an intelligent driving vehicle. Background Technology
[0002] In existing intelligent driving systems, camera-based obstacle detection methods typically track or fuse all detected obstacle targets directly, which involves a large amount of computation and is prone to introducing invalid information such as roadside stationary objects and distant false alarm targets into subsequent decision-making modules.
[0003] Specifically, some solutions use deep learning object detection models to directly identify and output single-frame images. However, these solutions have gradually revealed many significant technical shortcomings in real-world, complex driving scenarios. On the one hand, real-world driving environments are filled with numerous interfering factors: backlighting, low-light conditions at night, reflections from road surface water, dust kicked up by vehicles ahead, briefly swaying tree branches, and blurry road sign shadows in the distance. These factors can easily cause detection models to misclassify non-obstacle interfering targets as potential obstacles, generating a large number of low-confidence false detection results. On the other hand, existing technologies often employ a uniform filtering logic for all potential obstacles. They either discard all low-confidence detection results, easily missing small, partially occluded real obstacles in the distance, or perform complex feature matching and tracking calculations on all potential obstacles, significantly consuming the computing resources of the onboard computing platform and increasing the processing latency of the perception system, making it difficult to meet the low-latency and high-reliability operation requirements of intelligent driving systems.
[0004] Therefore, how to achieve stratified and differentiated screening of potential obstacles with different confidence levels in order to improve the accuracy of obstacle detection has become an urgent problem to be solved. Summary of the Invention
[0005] This invention provides a method for obstacle screening using a camera in an intelligent driving vehicle, the main purpose of which is to solve the problem of low obstacle detection accuracy.
[0006] Firstly, to achieve the above objectives, the present invention provides a method for obstacle screening using a camera in an intelligent driving vehicle, comprising: Acquire multiple consecutive target frame images captured by the onboard camera of an intelligent driving vehicle, identify multiple potential obstacles in all the target frame images, and obtain the recognition confidence level used to characterize each potential obstacle as a real obstacle; Based on the identification confidence level, the multiple potential obstacles are divided into a high-confidence obstacle set and a low-confidence obstacle set; Calculate the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images, and determine the motion state of each high-confidence obstacle based on the position change parameters; Using preset first screening conditions, a first obstacle in the set of high-confidence obstacles is determined based on the motion state, thus obtaining a first obstacle set; Extract the pixel size information and position information of each low-confidence obstacle in the target frame image from the set of low-confidence obstacles, and calculate the size confidence and position confidence of each low-confidence obstacle based on the pixel size information and the position information; Using preset second filtering conditions, based on the size confidence and location confidence, a second obstacle is determined from the set of low-confidence obstacles, thus obtaining a second obstacle set; The first obstacle set and the second obstacle set are merged to generate a candidate obstacle target set, and the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle is calculated. Obstacles whose collision threat index is greater than a preset collision threat threshold are identified as the target obstacle set.
[0007] Secondly, the present invention also provides a camera obstacle screening system for intelligent driving vehicles, the system comprising: An obstacle recognition module is used to acquire multiple consecutive target frame images captured by the on-board camera of an intelligent driving vehicle, identify multiple potential obstacles in all the target frame images, and obtain the recognition confidence level used to characterize each potential obstacle as a real obstacle. An obstacle classification module is used to classify multiple potential obstacles into a high-confidence obstacle set and a low-confidence obstacle set based on the recognition confidence level. The motion parameter analysis module is used to calculate the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images, and to determine the motion state of each high-confidence obstacle based on the position change parameters; The first screening module is used to determine the first obstacle in the set of high-confidence obstacles based on the motion state using preset first screening conditions, so as to obtain the first obstacle set. The confidence analysis module is used to extract the pixel size information and position information of each low-confidence obstacle in the low-confidence obstacle set in the target frame image, and to calculate the size confidence and position confidence of each low-confidence obstacle based on the pixel size information and the position information. The second filtering module is used to determine the second obstacle in the low-confidence obstacle set based on the size confidence and position confidence using preset second filtering conditions, so as to obtain the second obstacle set; The collision index calculation module is used to merge the first obstacle set and the second obstacle set to generate a candidate obstacle target set, and to calculate the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle. The obstacle determination module is used to determine the obstacle targets whose collision threat index is greater than a preset collision threat threshold as the target obstacle set.
[0008] In this embodiment of the invention, a confidence-driven dual-path differential screening mechanism is used to achieve accurate filtering and enhanced credibility of potential obstacles. Specifically, by separating high and low confidence levels, the high-overhead motion tracking and threat assessment are avoided for all targets. The secondary verification mechanism of the low-confidence set effectively retains difficult samples such as distant small targets and partially occluded targets, improving the system's detection recall rate in complex environments. At the same time, it realizes a hierarchical differential screening mechanism for potential obstacles with different confidence levels, thereby improving the obstacle detection accuracy. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart illustrating a method for obstacle screening using a camera in an intelligent driving vehicle, as provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of a camera obstacle screening system for intelligent driving vehicles according to an embodiment of the present invention; The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0011] To enable those skilled in the art to better understand the technical solutions of this disclosure, and to fully understand and implement the process of how this disclosure applies technical means to solve technical problems and achieve corresponding technical effects, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. The embodiments of this disclosure and the various features within them can be combined with each other without conflict, and the resulting technical solutions are all within the protection scope of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort should fall within the protection scope of this disclosure.
[0012] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0013] This application provides a method for obstacle screening using a camera in an intelligent driving vehicle. This method can be executed by software or hardware installed on a terminal device or server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0014] Reference Figure 1 The diagram shown is a flowchart illustrating a method for obstacle screening using a camera in an intelligent driving vehicle, according to an embodiment of the present invention. In this embodiment, the method includes: S1. Acquire multiple consecutive target frame images captured by the vehicle-mounted camera of the intelligent driving vehicle, identify multiple potential obstacles in all the target frame images, and obtain the recognition confidence level used to characterize each potential obstacle as a real obstacle.
[0015] In this embodiment of the invention, the target frame image refers to a single frame digital image extracted from a continuous video stream captured by an onboard camera of an intelligent driving vehicle at a specific time interval or triggering condition, used to identify and analyze objects in the road environment; the identification confidence score refers to the probability score output for each detected potential obstacle, used to characterize the degree of certainty that the identification result belongs to a real obstacle, and usually takes a value range of 0 to 1, where the higher the value, the greater the confidence that the target is a real obstacle.
[0016] In this embodiment of the invention, identifying multiple potential obstacles in all the target frame images includes: Each target frame image is subjected to image enhancement processing to obtain an enhanced image, and the probability value of each pixel in the enhanced image belonging to the obstacle category is calculated; All the probability values are collected to generate a probability map. Pixels with probability values greater than a preset obstacle probability threshold in the probability map are marked as obstacle pixels, and pixels with probability values less than or equal to the preset obstacle probability threshold in the probability map are marked as normal object pixels, thus obtaining a binarized obstacle mask image. Connectivity analysis is performed on the obstacle mask image to obtain all connected obstacle pixel regions, and the area of each obstacle pixel region is calculated. Based on the area, filter out obstacle pixel regions whose area is smaller than a preset area threshold to obtain a set of candidate obstacle regions; Extract the shape and texture features of each obstacle region in the candidate obstacle region set; A preset obstacle classifier is used to classify and identify each obstacle region based on its shape and texture features, thereby obtaining the obstacle category label and corresponding recognition confidence level for each obstacle region. Multiple potential obstacles in each obstacle region are identified based on the obstacle category labels.
[0017] In this embodiment of the invention, image enhancement processing is performed to address inherent defects in the target frame image, such as uneven illumination, blurred edges, and missing details. Specifically, by dynamically adjusting the pixel grayscale distribution and specifically enhancing high-frequency details in the target frame image, the feature differences between the obstacle and background areas are amplified without destroying the overall semantic structure of the image, thereby obtaining an enhanced image. Based on the pre-trained semantic segmentation feature mapping mechanism, the local texture, edge gradient, grayscale distribution, and other multi-dimensional visual features of the enhanced image are taken as input. Through the layer-by-layer abstraction of deep features and semantic association mapping, the confidence level of the obstacle category to which each pixel belongs is output. This process is completed based on the semantic matching logic of the feature space, realizing the quantification of the probability of category at the pixel level.
[0018] In detail, the category assignment probability values of all pixels are systematically collected according to the spatial arrangement rules of the original target frame image to form a probability distribution mapping carrier that is completely aligned with the spatial dimension of the original target frame image. This carrier completely retains the obstacle assignment confidence information of each pixel with spatial location as index, providing a globally unified probability distribution basis for subsequent pixel category classification.
[0019] Specifically, based on the preset semantic segmentation judgment threshold as the division boundary, all pixels in the probability map are divided into binary semantics. Pixels whose attribution confidence exceeds the judgment boundary are uniformly marked as obstacle attribute pixels, and the remaining pixels are marked as background attribute pixels. The binary semantic transformation of the image is completed through globally unified semantic judgment rules, generating a mask image that retains only the two types of semantic information: obstacle and background.
[0020] This process involves traversing all pixels marked as obstacle attributes within the binary mask image. Based on spatial adjacency rules between pixels, spatially connected pixels with the same attribute are aggregated into independent semantic regions. All independent obstacle pixel aggregation units are extracted, achieving structured region aggregation of discrete obstacle pixels. For each aggregated independent obstacle pixel region, the coverage area of all obstacle attribute pixels within that region is statistically analyzed. The spatial occupancy of each region is obtained through quantitative statistics of spatial coverage dimensions. This statistical result can intuitively reflect the actual spatial coverage of different obstacle regions, providing a core basis for subsequent noise region filtering.
[0021] Specifically, a screening operation is performed on all extracted obstacle pixel regions. Small regions whose spatial coverage does not meet the judgment criteria are directly identified as environmental noise or false targets generated by misidentification and are removed. Only valid regions that meet the spatial coverage requirements are retained. Finally, these regions are aggregated to form a set of candidate regions with the potential attributes of real obstacles, thus eliminating low-value misidentification interference items from the source.
[0022] Furthermore, for each independent obstacle region within the candidate region set, shape semantic information such as edge contour direction, overall topology, and boundary curvature distribution is extracted using contour morphology description technology. At the same time, deep texture semantic information such as grayscale fluctuation patterns and texture distribution patterns within the region is extracted using local texture statistical analysis technology. These two types of heterogeneous features are fused and encoded to form a high-dimensional feature vector that can fully characterize the visual attributes of the obstacle region.
[0023] Furthermore, based on the obstacle classification model that has been pre-trained under supervision, the high-dimensional feature vectors that integrate shape and texture information are extracted and used as input. Through the feature semantic matching and category mapping logic inside the model, the category classification of each candidate obstacle region is determined, and the confidence level quantification value corresponding to the category determination result is output simultaneously, providing classification semantic basis and confidence support for the final confirmation of potential obstacles.
[0024] Specifically, based on the semantic labels of the classifier output as the core judgment criteria, and combined with the spatial distribution association logic of the candidate regions, semantic verification and entity aggregation are carried out on each obstacle region that has been classified and labeled. Independent regions with unified semantic attributes are mapped to corresponding potential obstacle entities, completing the semantic transformation from pixel-level regions to entity-level obstacles, and finally realizing the complete identification and semantic labeling of all potential obstacles.
[0025] In this embodiment of the invention, continuous target frame images are acquired by an in-vehicle camera, and multi-frame collaborative perception is performed in combination with recognition confidence. This can effectively utilize temporal redundancy information to suppress accidental false detections and missed detections in single-frame detection, and significantly improve the accuracy of obstacle recognition.
[0026] S2. Based on the identification confidence level, the multiple potential obstacles are divided into a high-confidence obstacle set and a low-confidence obstacle set.
[0027] In this embodiment of the invention, the high-confidence obstacle set refers to the set of potential obstacles that are determined to have high confidence (usually higher than a preset threshold) during the confidence classification process. Obstacles in this set are considered more reliable detection results and tend to participate in subsequent decision-making responses with priority. The low-confidence obstacle set refers to the set of potential obstacles that are determined to have low confidence (usually lower than or equal to a preset threshold) during the confidence classification process. Obstacles in this set may have greater detection uncertainty and usually need to be further verified by combining more frame information and other sensor data to avoid false alarms leading to false braking or incorrect decisions.
[0028] In this embodiment of the invention, potential obstacles are dynamically divided into high and low confidence sets using a confidence threshold, enabling hierarchical and differentiated processing of perception results. The high confidence set can directly trigger robust tracking and response mechanisms, ensuring core requirements for driving safety. The low confidence set is guided to secondary verification or delayed decision-making processes, effectively avoiding unexpected actions such as false braking and emergency steering caused by single-frame false detections, thereby significantly reducing the false alarm rate and false alarm rate.
[0029] S3. Calculate the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images, and determine the motion state of each high-confidence obstacle based on the position change parameters.
[0030] In this embodiment of the invention, the position change parameter refers to the quantitative change index calculated by comparing the position coordinates of the same high-confidence obstacle in multiple consecutive target frame images. It typically includes the center point displacement rate, scale change rate, motion trajectory curvature, motion direction angle change, and position change statistical characteristics, etc., and is used to characterize the motion geometric characteristics of the obstacle on the image plane. The motion state refers to the qualitative or quantitative judgment made on the current dynamic attributes of the high-confidence obstacle based on the time series analysis results of the position change parameter. Common motion states include stationary, uniform motion, accelerated motion, decelerated motion, linear motion state, or curved motion state (turning or changing lanes), etc., and are used to describe the actual behavior pattern of the obstacle in space.
[0031] In this embodiment of the invention, calculating the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images includes: Determine the location information of the detection box of the high-confidence obstacle in multiple consecutive target frame images; The displacement vector of the center point of the high-confidence obstacle between two adjacent target frame images is calculated based on the center point coordinates of the detection box position information, and the displacement rate of the center point is calculated based on the time interval between the two adjacent target frame images. The scale change of the high-confidence obstacle between two adjacent target frame images is calculated based on the width and height of the detection box position information, and the scale change rate is calculated based on the time interval between the two adjacent target frame images. The curvature of the motion trajectory and the change in motion direction angle of the high-confidence obstacle in multiple consecutive target frame images are calculated based on the displacement vector of the center point and the scale change. Statistical features are extracted from the center point displacement rate, the scale change rate, the motion trajectory curvature, and the change in the motion direction angle to obtain the positional change statistical features of the high-confidence obstacle. The displacement rate of the center point, the rate of change of scale, the curvature of the motion trajectory, the change of the motion direction angle, and the statistical features of the position change are combined into a multidimensional position change parameter vector, which serves as the position change parameter of the high-confidence obstacle.
[0032] In this embodiment of the invention, obstacles considered to have high confidence are extracted from each frame of the image. Here, the "detection box" is a rectangular area that is accurately drawn around the obstacle using an object detection algorithm (such as a deep learning-based object detection network) to mark the specific position and size of the obstacle in the current image. From multiple consecutive target frame images, all detection box information of the same obstacle is extracted. This information includes the geometric attributes of the detection box (such as its center point coordinates, width, and height).
[0033] In detail, after obtaining the detection box positions in consecutive frames, the two-dimensional planar motion of the obstacle is quantified. Specifically, the detection boxes of the same obstacle in two adjacent frames are taken, and the position of their center points in the image coordinate system is calculated. A center point displacement vector is formed by pointing from the center point of the previous frame to the center point of the next frame. This vector describes the direction and distance of the obstacle's movement between the two frames. Simultaneously, since there is a fixed or known time interval between the two frames, the center point displacement rate of the obstacle is calculated by dividing the magnitude of this displacement vector by the time interval. This rate reflects how fast the obstacle moves on the current image plane.
[0034] Specifically, the visual size of an obstacle in an image changes due to its distance from the camera: it appears larger when closer to the camera and smaller when farther away. By reading the width and height values of the same obstacle detection box in two adjacent frames and comparing the changes in width and height between these two frames, a scale change can be obtained, which describes the degree to which the obstacle appears magnified or shrunk. Based on this, by dividing this scale change by the time interval between two frames, the scale change rate is obtained. This rate reflects the speed at which the obstacle approaches or moves away along the camera's optical axis (i.e., the depth direction), and is of great significance for determining whether the obstacle is moving towards the vehicle.
[0035] The curvature of the motion trajectory describes whether the obstacle's movement path is close to a straight line or a curve: if the direction of the displacement vector of the center point changes significantly over multiple consecutive frames, it indicates that the obstacle is turning or detouring, and the curvature value is high; conversely, if the direction of these vectors remains basically unchanged, it indicates that the movement is relatively straight, and the curvature value is low. Simultaneously, the change in motion direction angle is calculated, which is the difference in the direction angle between each two adjacent movements of the obstacle in a continuous frame sequence. This change reflects whether the obstacle is frequently adjusting its direction of travel, and is crucial for predicting its future path.
[0036] Furthermore, the calculation results of a single frame or a single adjacent frame often contain noise or random fluctuations. In order to obtain a more robust description of obstacle behavior, statistical analysis is performed on the above four indicators (i.e., center point displacement rate, scale change rate, motion trajectory curvature, and motion direction angle change) within a preset time window. That is, the statistics of these indicators over a period of time are calculated, such as the mean (reflecting the overall trend) and standard deviation (reflecting the degree of fluctuation), as well as the maximum value, minimum value, or range of change. These positional change statistical characteristics can reveal the stability, intensity, and overall pattern of obstacle movement behavior, providing more reliable data support for subsequent identification and decision-making.
[0037] Furthermore, the center point displacement rate, scale change rate, motion trajectory curvature, motion direction angle change amount calculated in each frame or time window, along with the position change statistical features extracted in step five, are combined into a multi-dimensional position change parameter vector, which serves as the position change parameter for high-confidence obstacles.
[0038] In this embodiment of the invention, determining the motion state of each high-confidence obstacle based on the position change parameters includes: Extract the center point displacement rate, scale change rate, motion trajectory curvature, and motion direction angle change from the position change parameters; If the displacement rate of the center point is less than or equal to a preset static displacement threshold, the high-confidence obstacle is determined to be a static obstacle; if the displacement rate of the center point is greater than the static displacement threshold, the high-confidence obstacle is determined to be a dynamic obstacle. When the high-confidence obstacle is a dynamic obstacle, the displacement rate of the center point is matched with multiple preset rate ranges, and the speed level of the high-confidence obstacle is determined based on the matching results. The consistency of the motion direction of the high-confidence obstacle is calculated based on the curvature of the motion trajectory and the change in the motion direction angle. If the consistency of the motion direction is greater than a preset consistency threshold, the high-confidence obstacle is determined to be in a straight-line motion state; if the consistency of the motion direction is less than or equal to the preset consistency threshold, the high-confidence obstacle is determined to be in a curved motion state. The scale change trend of the high-confidence obstacle is determined based on the comparison between the scale change rate and the preset scale change threshold. The motion state of the high-confidence obstacle is generated by combining the speed level, the linear motion state or the curvilinear motion state, and the scale change trend.
[0039] In this embodiment of the invention, the center point displacement rate, scale change rate, motion trajectory curvature, and motion direction angle change are extracted from the position change parameters. These indicators correspond to the translation speed of the obstacle on the image plane, the speed of approaching or moving away in the depth direction, the curvature of the path, and the adjustment frequency of the travel direction, respectively, and together constitute the original input for subsequent layer determination.
[0040] Specifically, the center point displacement rate is used as the most basic binary classification criterion, which compares the measured value with a preset threshold. The preset static displacement threshold represents a noise tolerance, which is the maximum displacement rate of an obstacle in an absolutely static state due to interference such as detection jitter and slight vibration. If the center point displacement rate of the current obstacle is lower than or equal to the threshold, it is determined that the obstacle does not have substantial spatial movement and is thus classified as a static obstacle (such as vehicles parked on the roadside, traffic cones, building walls, etc.). Conversely, if the displacement rate is higher than the threshold, it indicates that the obstacle has undergone significant displacement on the image plane and is classified as a dynamic obstacle (such as vehicles in motion, pedestrians walking, etc.).
[0041] When an obstacle is determined to be dynamic, its center point displacement rate is further quantified and graded. This involves mapping continuous displacement rate values to discrete speed levels (e.g., low, medium, high, or ultra-high speed) across multiple preset speed ranges. The measured displacement rate is compared one by one with the boundary values of these ranges to identify the range to which it belongs and assign a corresponding level label. Simultaneously, the motion direction consistency is calculated by combining the curvature of the motion trajectory and the change in the motion direction angle. The design philosophy of this index is: if the obstacle's motion direction is highly consistent across multiple frames (i.e., the change in direction angle is minimal and the trajectory curvature is close to zero), its motion direction consistency is high; conversely, if the direction changes frequently and the path is curved, the consistency is low. This index uses a certain aggregation calculation method to fuse information from the two dimensions into a single scalar to describe the stability of the obstacle's movement direction.
[0042] Specifically, if the consistency of an obstacle's movement direction is greater than this threshold, it indicates that its movement direction remains highly consistent within the time window, and its path approaches a straight line; in this case, it is determined to be in a straight-line motion state. If the consistency is less than or equal to this threshold, it indicates that the obstacle's movement direction has changed significantly, and its path exhibits curved, meandering, or serpentine characteristics; in this case, it is determined to be in a curvilinear motion state. This determination reflects the obstacle's driving intention and the complexity of its path, and has direct guiding significance for subsequent trajectory prediction and collision risk assessment.
[0043] Furthermore, the measured rate of scale change is compared with a preset scale change threshold to determine the obstacle's movement trend in the depth direction. A positive rate of scale change exceeding the threshold indicates that the obstacle is visually increasing in size, corresponding to its approaching the vehicle; a negative rate of scale change exceeding the threshold in absolute value indicates that the obstacle is visually decreasing in size, corresponding to its moving away from the vehicle; if the rate of scale change fluctuates within the threshold range, it is determined that the scale remains basically unchanged, meaning the relative depth relationship between the obstacle and the vehicle remains stable.
[0044] Furthermore, the judgment results from the above steps are fused from multiple dimensions to generate a structured and semantic description of the motion state. Specifically, the system combines three pieces of information—speed level (e.g., low speed, medium speed, high speed), motion form (linear motion or curvilinear motion), and scale change trend (approaching, moving away, or stable)—into a single overall label. For example, the final output might be "high-speed linear approaching motion," "medium-speed curvilinear moving away motion," or "low-speed linear stable motion," providing complete and interpretable input information for risk warning in autonomous driving systems.
[0045] In this embodiment of the invention, by tracking the positional change parameters (such as pixel displacement, direction vector, and inter-frame velocity) of high-confidence obstacles in consecutive multi-frame images, their motion state (such as stationary, uniform speed, acceleration, or turning) can be accurately characterized. This temporal motion analysis effectively utilizes the reliability of the high-confidence set and avoids interference from low-confidence noise, thereby obtaining a more stable and smooth motion trajectory estimate, significantly improving the accuracy of the system's understanding of dynamic obstacle behavior and situational awareness.
[0046] S4. Using the preset first screening conditions, determine the first obstacle in the set of high-confidence obstacles based on the motion state, and obtain the first obstacle set.
[0047] In this embodiment of the invention, the first screening condition refers to a preset judgment rule or standard for screening specific obstacles from a set of high-confidence obstacles. This condition is set based on the motion state of the obstacle (such as stationary, uniform speed, acceleration, etc.) and includes screening conditions such as speed level screening threshold, motion form preference type and allowable range of scale change, so as to distinguish obstacles with different behavior patterns and facilitate the execution of differentiated subsequent processing strategies.
[0048] In this embodiment of the invention, determining the first obstacle in the set of high-confidence obstacles based on the motion state using preset first screening conditions to obtain the first obstacle set includes: Extract the speed level screening threshold, motion morphology preference type, and scale variation allowable range from the preset first screening criteria; The speed level in the motion state of each high-confidence obstacle is compared with the speed level screening threshold to screen out high-confidence obstacles whose speed level meets the speed level screening threshold requirement, thus obtaining a first-level screening subset; From the first-level filtering subset, high-confidence obstacles whose motion patterns match the motion pattern preference type are selected according to the motion pattern preference type, to obtain the second-level filtering subset; From the second-level selection subset, high-confidence obstacles whose scale change trends fall within the allowable range of scale change are selected according to the allowable range of scale change, thus obtaining the third-level selection subset; All high-confidence obstacles in the third-level filtering subset are taken as first obstacles, and all first obstacles are aggregated to form a first obstacle set.
[0049] In this embodiment of the invention, three independent filtering dimension parameters are parsed from the pre-set first filtering conditions. The first is the speed level filtering threshold, which defines a speed level threshold—that is, only obstacles with speed levels higher than, equal to, or lower than a certain specific level are considered, the specific direction depending on the needs of the application scenario. The second is the motion pattern preference type, which specifies the obstacle motion patterns that the system expects to retain. It can choose to retain only obstacles with linear motion, only obstacles with curvilinear motion, or both. The third is the allowable range of scale variation, which defines an interval to limit the allowable range within which the obstacle's motion trend in the depth direction relative to the vehicle should fall.
[0050] Specifically, each obstacle in the high-confidence obstacle set is traversed, and its speed level in motion is extracted (this level has been categorized into discrete classes such as low speed, medium speed, high speed, or ultra-high speed in the previous process). This speed level is then compared with a preset speed level filtering threshold. The comparison logic includes three possibilities: the speed level must be greater than or equal to the threshold, less than or equal to the threshold, or exactly equal to a specified level. In the filtering results, all obstacles whose speed levels fall within the required range are retained, forming the first-level filtering subset to quickly eliminate obstacles whose speed characteristics do not conform to the task's focus. For example, in a scenario focused on high-speed rear-end collision warning, low-speed or stationary obstacles are excluded first.
[0051] Specifically, for each obstacle in the first-level filtering subset, its determined motion pattern attribute (linear motion or curvilinear motion) is checked and matched with a preset motion pattern preference type. If the preference type is set to "only retain linear motion", all obstacles determined to be in linear motion are retained, and those in curvilinear motion are eliminated; if the preference is set to "only retain curvilinear motion", the logic is reversed; if the preference is set to "no preference", all obstacles pass this level of filtering directly.
[0052] Specifically, for each obstacle in the second-level filtering subset, the scale change trend (approaching, moving away, or stable) in its motion state is extracted, and it is determined whether the trend falls within a preset allowable range of scale change. The allowable range of scale change is a discrete set, which can be set to allow only "approaching," only "moving away," only "stable," or a combination of multiple trends simultaneously. For example, in a forward collision warning scenario, only obstacles with a scale change trend of "approaching" may be allowed to pass the screening, because under the same conditions, only obstacles approaching the vehicle pose a potential collision threat, while obstacles that tend to move away or remain stable are eliminated at this level.
[0053] Among them, after the above three levels of screening, all high-confidence obstacles that successfully pass the three levels of screening are marked as first obstacles. All first obstacles are then gathered together to form a structured set of first obstacles.
[0054] In this embodiment of the invention, by using a preset first screening condition and motion state as the criterion, the first obstacle is accurately extracted from the high-confidence set, achieving contextualized filtering and hierarchical focusing of the perception results. This mechanism can effectively filter out moving objects that do not directly affect the current driving behavior (such as distant vehicles traveling in the same direction), while prioritizing the retention of key targets that require immediate response (such as stationary or rapidly decelerating vehicles in front), thereby significantly reducing the ineffective computational load of downstream risk decisions and effectively improving the driving safety of the intelligent driving system in complex traffic flow.
[0055] S5. Extract the pixel size information and position information of each low-confidence obstacle in the target frame image from the low-confidence obstacle set, and calculate the size confidence and position confidence of each low-confidence obstacle based on the pixel size information and the position information.
[0056] In this embodiment of the invention, pixel size information refers to the size of the pixel region occupied by a low-confidence obstacle in the target frame image, usually represented by the width and height of the bounding box (in pixels) or area, used to characterize the visual scale of the obstacle on the image plane. Location information refers to the spatial positioning data of the low-confidence obstacle in the target frame image, usually represented by the image coordinates of the center point of the bounding box or the corner coordinates of a rectangular area, used to characterize the spatial distribution position of the obstacle on the image plane. Size confidence is a reliability metric calculated by comparing the pixel size information (such as width, height, and area) of the low-confidence obstacle with a preset prior size distribution or typical size range, used to characterize whether the size estimate of the obstacle is reliable; a higher value indicates that the scale of the obstacle is more in line with expectations. Location confidence is a reliability metric calculated by comparing the location information (such as image coordinates and distance from the image edge) of the low-confidence obstacle with a preset prior location distribution or drivable area range, used to characterize whether the location estimate of the obstacle is reliable; a higher value indicates that the spatial positioning of the obstacle is more in line with expectations.
[0057] In this embodiment of the invention, extracting the pixel size information and position information of each low-confidence obstacle in the low-confidence obstacle set in the target frame image includes: Obtain the target frame image corresponding to the low-confidence obstacle, and locate the bounding box coordinate information of the low-confidence obstacle from the target frame image; The pixel width and pixel height of the low-confidence obstacle in the target frame image are calculated based on the bounding box coordinate information, and the pixel width and pixel height are used as the initial pixel size information of the low-confidence obstacle. The bounding box coordinate information is corrected for coordinate precision to obtain the target bounding box coordinate information; The corrected pixel width and corrected pixel height of the low-confidence obstacle in the target frame image are calculated based on the target bounding box coordinate information, and the corrected pixel width and corrected pixel height are used as the pixel size information of the low-confidence obstacle. The center point coordinates of the low-confidence obstacle are calculated based on the target bounding box coordinate information, and the center point coordinates are used as the position information of the low-confidence obstacle.
[0058] In this embodiment of the invention, the target image frame corresponding to the current low-confidence obstacle is located. On this frame, the result output by the target detection algorithm is called to obtain the bounding box coordinate information when the obstacle is detected. Here, the bounding box is a rectangular area, which is usually defined by the coordinates of its upper left and lower right corners in the image coordinate system, or by the combination of the center coordinates and the width and height.
[0059] Specifically, using the bounding box coordinate information obtained in the first step, the pixel width and pixel height of the obstacle are derived through simple geometric calculations. The pixel width is determined by the horizontal span of the bounding box, and the pixel height is determined by the vertical span of the bounding box. These two values together constitute the initial pixel size information of the obstacle in the current frame image.
[0060] In detail, to address potential jitter, offset, or scale inaccuracies in the initial bounding box coordinates, a coordinate accuracy correction process is introduced. Methods employed for coordinate accuracy correction typically include, but are not limited to, using temporal filtering techniques. By combining the bounding box position of the obstacle in adjacent frames, the coordinates of the current frame are smoothly corrected to eliminate random jitter caused by detection instability. After correction, a new set of coordinate data, i.e., the target bounding box coordinate information, is obtained, with a significantly improved positioning accuracy compared to the initial data.
[0061] Furthermore, based on the target bounding box coordinate information, the pixel width and pixel height of the obstacle in the current frame image are recalculated. These corrected values are formally defined as the pixel size information of the low-confidence obstacle and stored as the final size feature of the obstacle in the current frame. After completing the coordinate correction and size calculation, the center point coordinates of the low-confidence obstacle are calculated using the target bounding box coordinate information. The center point coordinates are usually determined by the midpoint between the left and right boundaries of the bounding box for the horizontal position and by the midpoint between the upper and lower boundaries for the vertical position, resulting in a two-dimensional coordinate point. This center point coordinate is defined as the position information of the low-confidence obstacle.
[0062] In this embodiment of the invention, calculating the size confidence and location confidence of each low-confidence obstacle based on the pixel size information and the location information includes: Obtain a preset reference size range, which includes a reference width range and a reference height range; The width confidence component is calculated based on the relative position of the pixel width in the pixel size information within the reference width range, and the height confidence component is calculated based on the relative position of the pixel height in the pixel size information within the reference height range. The width confidence component and the height confidence component are weighted and fused according to a preset first fusion weight to obtain the size confidence of the low-confidence obstacle; Obtain a preset reference position area, and determine whether the center point coordinates in the position information fall within the reference position area; If the center point coordinates in the location information fall within the reference location area, then the first location confidence component is calculated based on the offset distance between the center point coordinates and the center point of the reference location area. Obtain the position offset sequence of the low-confidence obstacle in multiple consecutive target frame images, calculate the position stability value of the low-confidence obstacle based on the position offset sequence, and use the position stability value as the second position confidence component; The first position confidence component and the second position confidence component are weighted and fused according to the preset second fusion weight to obtain the position confidence of the low-confidence obstacle.
[0063] In this embodiment of the invention, a reference size range is read from preset configuration parameters. This range is not a single numerical value, but rather includes two independent dimensions—a reference width range and a reference height range. These two ranges define the normal pixel width and pixel height range that one or more typical obstacles (such as cars, SUVs, pedestrians, bicycles, etc.) should have in the image. The upper and lower boundaries of the reference range are typically derived from statistical results of a large amount of historical data, or from the mapping relationship between sensor calibration parameters and typical physical dimensions.
[0064] Specifically, the pixel width of the current low-confidence obstacle is compared with a reference width range, and a width confidence component is generated based on the relative position of the width value within the range. Specifically, if the pixel width falls exactly in the middle of the reference width range, it indicates that the visual width of the obstacle matches the common width and height of typical obstacles of this type, and the width confidence component is assigned a higher value. If the pixel width deviates from the boundary of the range or even exceeds the range, it indicates that the visual width of the obstacle deviates from the typical distribution, and the width confidence component is correspondingly reduced. The same logic is independently applied to the comparison of pixel height with a reference height range to generate a height confidence component.
[0065] In detail, the width confidence component and the height confidence component are weighted and fused according to a preset first fusion weight. For example, in some application scenarios, width may be more discriminative than height (such as when distinguishing between vehicles and pedestrians, where the difference in width is more significant), so the width confidence component is given a higher fusion weight; in other scenarios, the weights of the two components may be equal. The result of the weighted fusion is a single numerical value that systematically integrates the confidence information of both the width and height dimensions, and is ultimately defined as the size confidence of the low-confidence obstacle.
[0066] The process involves reading a preset reference position region from the configuration parameters. This region is typically defined within a specific sub-region of the image, such as the area directly in front of the lane, the projection area of the vehicle's trajectory, or the region of interest. The reference position region represents the spatial range of focus. The determination of whether the center point coordinates of the current low-confidence obstacle fall within this reference position region will determine the calculation path for the subsequent first position confidence component.
[0067] Specifically, if the center point coordinates of the current low-confidence obstacle do indeed fall within the reference position area, the offset distance between this center point and the center point of the reference position area is further calculated. The smaller the offset distance, the closer the spatial position of the obstacle is to the center of the area of interest, the more typical its position, and the higher the value of the first position confidence component. Conversely, the larger the offset distance, the closer the obstacle is to the edge of the area of interest, the more its position deviates from the typical scenario, and the lower the first position confidence component is accordingly. If the center point coordinates do not fall within the reference position area, this component will be assigned a lower default value or be set to zero.
[0068] This process involves extracting the positional offset sequence of the low-confidence obstacle across multiple consecutive target frame images. This sequence represents the displacement of the obstacle's center point between adjacent frames over time. A positional stability value is then calculated for this sequence, analyzing whether the obstacle's movement between consecutive frames is smooth and consistent. If the magnitude and direction of the positional offset remain stable across frames (e.g., without significant sudden jumps or reversals), the obstacle's positional data exhibits high temporal consistency and reliability, resulting in a high positional stability value. Conversely, if the positional offset fluctuates wildly or its direction changes frequently, it indicates significant jitter or uncertainty in the detection results, leading to a low positional stability value. This stability value is defined as the second positional confidence component.
[0069] Furthermore, according to the preset second fusion weight, the first position confidence component (reflecting the typicality of spatial position) and the second position confidence component (reflecting the stability of temporal position) are weighted and fused to obtain the position confidence of the low confidence obstacle.
[0070] In this embodiment of the invention, for obstacles with low confidence, by extracting their pixel size and location information and calculating size confidence and location confidence, a multi-dimensional confidence assessment of weakly detected targets is achieved. This provides auxiliary judgment criteria from two independent perspectives—physical scale and spatial distribution—when the recognition confidence is insufficient. Thus, it can significantly improve the retention rate of potential effective targets in low-confidence sets without increasing hardware costs, while effectively filtering out geometrically inconsistent false detections and enhancing the comprehensive processing capability under recognition uncertainty.
[0071] S6. Using the preset second screening conditions, determine the second obstacle in the low-confidence obstacle set based on the size confidence and position confidence, and obtain the second obstacle set.
[0072] In this embodiment of the invention, the step of determining the second obstacle in the low-confidence obstacle set based on the size confidence and location confidence using preset second screening conditions to obtain the second obstacle set includes: Obtain a preset second filtering condition, which includes at least a size confidence lower limit threshold and a location confidence lower limit threshold; Based on the second screening criteria, low-confidence obstacles with a size confidence level greater than or equal to the lower limit threshold of the size confidence level are selected to obtain a first confidence screening subset; From the first confidence filtering subset, low-confidence obstacles whose location confidence is greater than or equal to the lower limit threshold of location confidence are filtered out according to the second filtering condition to obtain the second confidence filtering subset; For each low-confidence obstacle in the second confidence screening subset, calculate the weighted average of the size confidence and the location confidence; Low-confidence obstacles whose weighted average value is greater than or equal to a preset confidence threshold are designated as second obstacles, and all second obstacles are aggregated to form a second obstacle set.
[0073] In this embodiment of the invention, a second filtering condition is read from preset configuration parameters. This filtering condition includes at least two core threshold parameters: a size confidence threshold and a location confidence threshold. The size confidence threshold defines a minimum acceptable size confidence level—only obstacles whose size characteristics are sufficiently close to the expected distribution are likely to pass the filtering; the location confidence threshold defines a minimum threshold for spatial location confidence—only obstacles whose spatial location is sufficiently typical and whose temporal location is sufficiently stable are likely to be retained.
[0074] In detail, each obstacle in the low-confidence obstacle set is traversed, and its size confidence value calculated in the previous process is extracted. This size confidence value is then compared with a preset size confidence lower limit threshold. Obstacles with a size confidence value greater than or equal to the threshold are considered to have reached a basically acceptable level of matching between their pixel width and height and the reference size range. They are considered to have passed the size dimension confidence test and are retained to form the first confidence screening subset. Obstacles with a size confidence value lower than the threshold mean that their visual size deviates significantly from the characteristic distribution of typical obstacles, and their size information is unreliable. They are directly eliminated.
[0075] In detail, for each obstacle in the first confidence screening subset, its position confidence value is further extracted, and the position confidence value is compared with a preset lower limit threshold for position confidence. Obstacles with a position confidence value greater than or equal to the threshold indicate that their spatial position not only meets the position typicality requirements of the area of interest (first component) but also shows sufficient motion stability in the time series (second component). They are considered to have passed the confidence test of the position dimension and are retained to form the second confidence screening subset. Obstacles with a position confidence value lower than the threshold will be eliminated even if their size confidence value meets the standard because their position information is unreliable.
[0076] For each obstacle in the second confidence subset that has passed the dual-threshold screening, a weighted average of its size confidence and location confidence is further calculated. This weighting reflects the system's preference for the relative importance of size and location information in the overall confidence assessment. The calculation of the weighted average merges the confidence from both dimensions into a unified comprehensive score. This score more comprehensively measures the overall credibility of the obstacle, avoiding potential biases that might arise from making a final judgment based on only a single dimension.
[0077] Furthermore, the weighted average is compared with a preset confidence threshold: obstacles with a weighted average greater than or equal to the threshold are determined to have sufficient confidence at the overall level and are thus determined to be second obstacles; obstacles with a weighted average lower than the threshold, even if they have passed the first two steps of screening, are still eliminated due to insufficient overall confidence. All second obstacles that have successfully passed this final determination are gathered together to form a second obstacle set.
[0078] In this embodiment of the invention, by using a preset second screening condition, second obstacles are extracted from the low-confidence set based on size confidence and location confidence, thereby achieving secondary verification and effective recall of weakly detected targets. This enables hierarchical and differentiated screening of potential obstacles with different confidence levels, thereby improving the accuracy of obstacle detection.
[0079] S7. Merge the first obstacle set and the second obstacle set to generate a candidate obstacle target set, and calculate the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle.
[0080] In this embodiment of the invention, the collision threat index refers to a comprehensive assessment parameter calculated for each obstacle target in the candidate obstacle target set to quantify the degree of potential collision risk between it and the intelligent driving vehicle. This index is usually calculated by combining multi-dimensional information such as the obstacle's motion state, relative distance, relative speed, relative acceleration, target size, and predicted trajectory through a preset collision risk assessment model, and is used to characterize the urgency of the threat posed by the target to the vehicle.
[0081] In this embodiment of the invention, calculating the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle includes: The current self-driving vehicle motion state parameters and the obstacle motion state parameters of each obstacle target in the candidate obstacle target set are obtained. The self-driving vehicle motion state parameters include at least the self-driving vehicle speed and the self-driving vehicle direction angle, and the obstacle motion state parameters include at least the obstacle speed and the obstacle direction angle. The relative driving direction angle between the intelligent driving vehicle and the obstacle target is calculated based on the driving direction angle of the vehicle and the driving direction angle of the obstacle, and the orientation area of the obstacle target relative to the intelligent driving vehicle is determined based on the relative driving direction angle. The relative velocity vector between the intelligent driving vehicle and the obstacle target is calculated based on the vehicle speed and the obstacle speed, and the relative approach rate between the intelligent driving vehicle and the obstacle target is calculated based on the relative velocity vector. The current spatial distance between the intelligent driving vehicle and the obstacle target is obtained, and the collision time between the intelligent driving vehicle and the obstacle target is calculated based on the current spatial distance and the relative approach rate; The azimuth threat weight coefficient is determined based on the azimuth region, the rate threat weight coefficient is determined based on the relative approach rate, and the time threat weight coefficient is determined based on the collision time. The directional threat weight coefficient, the rate threat weight coefficient, and the time threat weight coefficient are weighted and fused according to the preset threat fusion weight, and the weighted fusion result is normalized to obtain the collision threat index between the obstacle target and the intelligent driving vehicle.
[0082] In this embodiment of the invention, the current motion state parameters of the vehicle are obtained from the perception and positioning module of the intelligent driving vehicle, including at least the instantaneous speed and the current driving direction angle of the vehicle. Simultaneously, the motion state parameters of each obstacle are read one by one from the candidate obstacle target set, also including at least the speed and driving direction angle of the obstacle.
[0083] Specifically, the driving direction angle of the vehicle is compared and analyzed with that of the obstacle to calculate the angle of deviation between them, i.e., the relative driving direction angle. The size of this angle directly reflects the degree of convergence of the driving and obstacle's directions of motion: the smaller the angle, the more likely they are moving in the same direction; an angle close to opposite directions indicates they are moving towards each other; and an angle close to perpendicular indicates a possibility of crossing. Based on this, according to the range of values for this angle, the obstacle's position relative to the vehicle is divided into different directional zones, such as a forward-facing zone, a frontal facing zone, a lateral crossing zone, or a rear-facing zone. Different directional zones correspond to different potential collision patterns and hazard levels.
[0084] Specifically, the vehicle's speed and the obstacle's speed are vector-combined to obtain the relative speed vector between them. This vector describes the obstacle's motion state in the vehicle's reference frame, that is, the speed and direction of the obstacle's motion if the vehicle is taken as a stationary reference. From this relative speed vector, the component along the line connecting the vehicle and the obstacle is further extracted, which is the relative approach speed. When the relative approach speed is positive, it indicates that the obstacle is approaching the vehicle along the line connecting them. When the relative approach speed is negative, it indicates that the obstacle is moving away from the vehicle.
[0085] Specifically, the spatial distance between the vehicle and the obstacle at the current moment is obtained (usually provided by sensors such as millimeter-wave radar, lidar, or stereo vision). Combined with the relative approach rate, the time required for the vehicle to make contact with the obstacle while maintaining its current motion state is inferred, i.e., the collision time. The shorter the collision time, the less margin is left for the vehicle to respond and avoid, and the higher the corresponding threat level.
[0086] Furthermore, based on the location region, a preset mapping relationship is queried to obtain the location threat weight coefficient. For example, obstacles directly in front of the vehicle are usually assigned a higher location threat weight, while obstacles behind it have a lower weight. Based on the relative approach rate, a preset rate threat weight mapping table is queried; the higher the relative approach rate, the higher the rate threat weight coefficient. Based on the collision time, a preset time threat weight mapping table is queried; the shorter the collision time, the higher the time threat weight coefficient.
[0087] Furthermore, according to the preset threat fusion weights, the directional threat weight coefficient, the rate threat weight coefficient, and the time threat weight coefficient are weighted and fused together to obtain a comprehensive original threat score. The original threat score is then normalized and mapped to a uniform numerical range (such as between zero and one or between zero and one hundred) to finally obtain the collision threat index between the obstacle target and the intelligent driving vehicle.
[0088] In this embodiment of the invention, by merging the first obstacle set and the second obstacle set to generate a candidate obstacle target set, the information complementarity and fusion of high-confidence reliable targets and low-confidence recall targets are achieved, avoiding the need to process all targets with the same level of complexity. Furthermore, it enables hierarchical and differentiated screening of potential obstacles with different levels of confidence, thereby improving the obstacle detection accuracy.
[0089] S8. Determine whether the collision threat index is greater than the preset collision threat threshold.
[0090] S9. Obstacles whose collision threat index is greater than a preset collision threat threshold are identified as the target obstacle set.
[0091] S10. Obstacles whose collision threat index is less than or equal to a preset collision threat threshold are identified as normal objects.
[0092] In this embodiment of the invention, a collision threat level threshold is preset. This threshold represents the highest permissible level of safety risk. Any threat level index higher than this level is considered an unacceptable dangerous state, while any level lower than or equal to this level is considered to be within the safe tolerance range under the current conditions.
[0093] Specifically, the process iterates through each obstacle in the candidate obstacle target set, comparing its collision threat index calculated in the preceding process with a preset threshold. Obstacles with a collision threat index greater than the preset threshold are classified as dangerous targets posing an actual threat and added to the target obstacle set. Obstacles with a collision threat index less than or equal to the preset threshold are classified as normal targets that do not pose an imminent threat in the current state and added to the normal object set. This binary classification enables differentiated processing of all obstacles within the perception range.
[0094] In this embodiment of the invention, a differentiated processing mechanism is achieved by introducing a preset collision threat threshold to dynamically divide and determine the candidate obstacle target set. The threshold comparison judgment in S8, as the core switch for decision-making, can convert continuous threat quantification values into discrete decision branches, providing a clear and interpretable basis for risk classification. S9 categorizes high-risk targets into the target obstacle set, enabling key safety response mechanisms such as emergency collision avoidance, braking, or steering to accurately focus on truly threatening targets, effectively avoiding decision-making chaos or over-response problems caused by treating all targets equally. S10 categorizes low-risk targets into the normal object set, allowing standard driving behaviors such as normal cruise and adaptive following to be stably maintained, unaffected by irrelevant or distant targets.
[0095] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0096] like Figure 2 The diagram shown is a functional block diagram of a camera obstacle screening system for intelligent driving vehicles provided in an embodiment of the present invention.
[0097] This disclosure provides an obstacle screening system for cameras used in intelligent driving vehicles, which corresponds one-to-one with the obstacle screening method for cameras used in intelligent driving vehicles described in the previous embodiment. For example... Figure 2 As shown, the intelligent driving vehicle camera obstacle screening system 100 includes an obstacle recognition module 101, an obstacle classification module 102, a motion parameter analysis module 103, a first screening module 104, a confidence analysis module 105, a second screening module 106, a collision index calculation module 107, and an obstacle determination module 108. Detailed descriptions of each functional module are as follows: Obstacle recognition 101 is used to acquire multiple consecutive target frame images captured by the on-board camera of the intelligent driving vehicle, identify multiple potential obstacles in all the target frame images, and obtain the recognition confidence level used to characterize each potential obstacle as a real obstacle. The obstacle classification module 102 is used to classify a plurality of potential obstacles into a high-confidence obstacle set and a low-confidence obstacle set according to the recognition confidence level; The motion parameter analysis module 103 is used to calculate the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images, and to determine the motion state of each high-confidence obstacle based on the position change parameters. The first screening module 104 is used to determine the first obstacle in the set of high-confidence obstacles based on the motion state using preset first screening conditions, and to obtain the first obstacle set. The confidence analysis module 105 is used to extract the pixel size information and position information of each low-confidence obstacle in the set of low-confidence obstacles in the target frame image, and to calculate the size confidence and position confidence of each low-confidence obstacle based on the pixel size information and the position information. The second filtering module 106 is used to determine the second obstacle in the low-confidence obstacle set based on the size confidence and position confidence using preset second filtering conditions, so as to obtain the second obstacle set; The collision index calculation module 107 is used to merge the first obstacle set and the second obstacle set to generate a candidate obstacle target set, and calculate the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle. The obstacle determination module 108 is used to determine the obstacle targets whose collision threat index is greater than the preset collision threat threshold as the target obstacle set.
[0098] In this invention, the specific limitations of a camera obstacle screening system for intelligent driving vehicles can be found in the above-described limitations of a camera obstacle screening method for intelligent driving vehicles, and will not be repeated here. Each module in the aforementioned camera obstacle screening system for intelligent driving vehicles can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0099] In the embodiments provided by this invention, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0100] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0101] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0102] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented. Any references to memory, storage, databases, or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0104] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0105] In the embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0106] It should be noted that, in this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element limited by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0107] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for obstacle screening using a camera in an intelligent driving vehicle, characterized in that, The method includes: Acquire multiple consecutive target frame images captured by the onboard camera of an intelligent driving vehicle, identify multiple potential obstacles in all the target frame images, and obtain the recognition confidence level used to characterize each potential obstacle as a real obstacle; Based on the identification confidence level, the multiple potential obstacles are divided into a high-confidence obstacle set and a low-confidence obstacle set; Calculate the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images, and determine the motion state of each high-confidence obstacle based on the position change parameters; Using preset first screening conditions, a first obstacle in the set of high-confidence obstacles is determined based on the motion state, thus obtaining a first obstacle set; Extract the pixel size information and position information of each low-confidence obstacle in the target frame image from the set of low-confidence obstacles, and calculate the size confidence and position confidence of each low-confidence obstacle based on the pixel size information and the position information; Using preset second filtering conditions, based on the size confidence and location confidence, a second obstacle is determined from the set of low-confidence obstacles, thus obtaining a second obstacle set; The first obstacle set and the second obstacle set are merged to generate a candidate obstacle target set, and the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle is calculated. Obstacles whose collision threat index is greater than a preset collision threat threshold are identified as the target obstacle set.
2. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The identification of multiple potential obstacles in all the target frame images includes: Each target frame image is subjected to image enhancement processing to obtain an enhanced image, and the probability value of each pixel in the enhanced image belonging to the obstacle category is calculated; All the probability values are collected to generate a probability map. Pixels with probability values greater than a preset obstacle probability threshold in the probability map are marked as obstacle pixels, and pixels with probability values less than or equal to the preset obstacle probability threshold in the probability map are marked as normal object pixels, thus obtaining a binarized obstacle mask image. Connectivity analysis is performed on the obstacle mask image to obtain all connected obstacle pixel regions, and the area of each obstacle pixel region is calculated. Based on the area, filter out obstacle pixel regions whose area is smaller than a preset area threshold to obtain a set of candidate obstacle regions; Extract the shape and texture features of each obstacle region in the candidate obstacle region set; A preset obstacle classifier is used to classify and identify each obstacle region based on its shape and texture features, thereby obtaining the obstacle category label and corresponding recognition confidence level for each obstacle region. Multiple potential obstacles in each obstacle region are identified based on the obstacle category labels.
3. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The calculation of the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images includes: Determine the location information of the detection box of the high-confidence obstacle in multiple consecutive target frame images; The displacement vector of the center point of the high-confidence obstacle between two adjacent target frame images is calculated based on the center point coordinates of the detection box position information, and the displacement rate of the center point is calculated based on the time interval between the two adjacent target frame images. The scale change of the high-confidence obstacle between two adjacent target frame images is calculated based on the width and height of the detection box position information, and the scale change rate is calculated based on the time interval between the two adjacent target frame images. The curvature of the motion trajectory and the change in motion direction angle of the high-confidence obstacle in multiple consecutive target frame images are calculated based on the displacement vector of the center point and the scale change. Statistical features are extracted from the center point displacement rate, the scale change rate, the motion trajectory curvature, and the change in the motion direction angle to obtain the positional change statistical features of the high-confidence obstacle. The displacement rate of the center point, the rate of change of scale, the curvature of the motion trajectory, the change of the motion direction angle, and the statistical features of the position change are combined into a multidimensional position change parameter vector, which serves as the position change parameter of the high-confidence obstacle.
4. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, Determining the motion state of each high-confidence obstacle based on the position change parameters includes: Extract the center point displacement rate, scale change rate, motion trajectory curvature, and motion direction angle change from the position change parameters; If the displacement rate of the center point is less than or equal to a preset static displacement threshold, the high-confidence obstacle is determined to be a static obstacle; if the displacement rate of the center point is greater than the static displacement threshold, the high-confidence obstacle is determined to be a dynamic obstacle. When the high-confidence obstacle is a dynamic obstacle, the displacement rate of the center point is matched with multiple preset rate ranges, and the speed level of the high-confidence obstacle is determined based on the matching results. The consistency of the motion direction of the high-confidence obstacle is calculated based on the curvature of the motion trajectory and the change in the motion direction angle. If the consistency of the motion direction is greater than a preset consistency threshold, the high-confidence obstacle is determined to be in a straight-line motion state; if the consistency of the motion direction is less than or equal to the preset consistency threshold, the high-confidence obstacle is determined to be in a curved motion state. The scale change trend of the high-confidence obstacle is determined based on the comparison between the scale change rate and the preset scale change threshold. The motion state of the high-confidence obstacle is generated by combining the speed level, the linear motion state or the curvilinear motion state, and the scale change trend.
5. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The first obstacle in the set of high-confidence obstacles is determined based on the motion state using preset first screening conditions, resulting in the first obstacle set, including: Extract the speed level screening threshold, motion morphology preference type, and scale variation allowable range from the preset first screening criteria; The speed level in the motion state of each high-confidence obstacle is compared with the speed level screening threshold to screen out high-confidence obstacles whose speed level meets the speed level screening threshold requirement, thus obtaining a first-level screening subset; From the first-level filtering subset, high-confidence obstacles whose motion patterns match the motion pattern preference type are selected according to the motion pattern preference type, to obtain the second-level filtering subset; From the second-level selection subset, high-confidence obstacles whose scale change trends fall within the allowable scale change range are selected according to the allowable scale change range, thus obtaining the third-level selection subset; All high-confidence obstacles in the third-level filtering subset are taken as first obstacles, and all first obstacles are aggregated to form a first obstacle set.
6. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The step of extracting the pixel size and position information of each low-confidence obstacle in the set of low-confidence obstacles in the target frame image includes: Obtain the target frame image corresponding to the low-confidence obstacle, and locate the bounding box coordinate information of the low-confidence obstacle from the target frame image; The pixel width and pixel height of the low-confidence obstacle in the target frame image are calculated based on the bounding box coordinate information, and the pixel width and pixel height are used as the initial pixel size information of the low-confidence obstacle. The bounding box coordinate information is corrected for coordinate precision to obtain the target bounding box coordinate information; The corrected pixel width and corrected pixel height of the low-confidence obstacle in the target frame image are calculated based on the target bounding box coordinate information, and the corrected pixel width and corrected pixel height are used as the pixel size information of the low-confidence obstacle. The center point coordinates of the low-confidence obstacle are calculated based on the target bounding box coordinate information, and the center point coordinates are used as the position information of the low-confidence obstacle.
7. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The step of calculating the size confidence and location confidence of each low-confidence obstacle based on the pixel size information and the location information includes: Obtain a preset reference size range, which includes a reference width range and a reference height range; The width confidence component is calculated based on the relative position of the pixel width in the pixel size information within the reference width range, and the height confidence component is calculated based on the relative position of the pixel height in the pixel size information within the reference height range. The width confidence component and the height confidence component are weighted and fused according to a preset first fusion weight to obtain the size confidence of the low-confidence obstacle; Obtain a preset reference position area, and determine whether the center point coordinates in the position information fall within the reference position area; If the center point coordinates in the location information fall within the reference location area, then the first location confidence component is calculated based on the offset distance between the center point coordinates and the center point of the reference location area. Obtain the position offset sequence of the low-confidence obstacle in multiple consecutive target frame images, calculate the position stability value of the low-confidence obstacle based on the position offset sequence, and use the position stability value as the second position confidence component; The first position confidence component and the second position confidence component are weighted and fused according to the preset second fusion weight to obtain the position confidence of the low-confidence obstacle.
8. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The second obstacle set is obtained by using preset second screening conditions based on the size confidence and location confidence to determine the second obstacle in the low-confidence obstacle set, including: Obtain a preset second filtering condition, which includes at least a size confidence lower limit threshold and a location confidence lower limit threshold; Based on the second screening criteria, low-confidence obstacles with a size confidence level greater than or equal to the lower limit threshold of the size confidence level are selected to obtain a first confidence screening subset; From the first confidence filtering subset, low-confidence obstacles whose location confidence is greater than or equal to the lower limit threshold of location confidence are filtered out according to the second filtering condition to obtain the second confidence filtering subset; For each low-confidence obstacle in the second confidence screening subset, calculate the weighted average of the size confidence and the location confidence; Low-confidence obstacles whose weighted average value is greater than or equal to a preset confidence threshold are designated as second obstacles, and all second obstacles are aggregated to form a second obstacle set.
9. The obstacle screening method for intelligent driving vehicles using cameras as described in claim 1, characterized in that, The calculation of the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle includes: The current self-driving vehicle motion state parameters and the obstacle motion state parameters of each obstacle target in the candidate obstacle target set are obtained. The self-driving vehicle motion state parameters include at least the self-driving vehicle speed and the self-driving vehicle direction angle, and the obstacle motion state parameters include at least the obstacle speed and the obstacle direction angle. The relative driving direction angle between the intelligent driving vehicle and the obstacle target is calculated based on the driving direction angle of the vehicle and the driving direction angle of the obstacle, and the orientation area of the obstacle target relative to the intelligent driving vehicle is determined based on the relative driving direction angle. The relative velocity vector between the intelligent driving vehicle and the obstacle target is calculated based on the vehicle speed and the obstacle speed, and the relative approach rate between the intelligent driving vehicle and the obstacle target is calculated based on the relative velocity vector. The current spatial distance between the intelligent driving vehicle and the obstacle target is obtained, and the collision time between the intelligent driving vehicle and the obstacle target is calculated based on the current spatial distance and the relative approach rate; The azimuth threat weight coefficient is determined based on the azimuth region, the rate threat weight coefficient is determined based on the relative approach rate, and the time threat weight coefficient is determined based on the collision time. The directional threat weight coefficient, the rate threat weight coefficient, and the time threat weight coefficient are weighted and fused according to the preset threat fusion weight, and the weighted fusion result is normalized to obtain the collision threat index between the obstacle target and the intelligent driving vehicle.
10. A camera obstacle screening system for intelligent driving vehicles, characterized in that, The system includes: An obstacle recognition module is used to acquire multiple consecutive target frame images captured by the on-board camera of an intelligent driving vehicle, identify multiple potential obstacles in all the target frame images, and obtain the recognition confidence level used to characterize each potential obstacle as a real obstacle. An obstacle classification module is used to classify multiple potential obstacles into a high-confidence obstacle set and a low-confidence obstacle set based on the recognition confidence level. The motion parameter analysis module is used to calculate the position change parameters of each high-confidence obstacle in the set of high-confidence obstacles in multiple consecutive target frame images, and to determine the motion state of each high-confidence obstacle based on the position change parameters; The first screening module is used to determine the first obstacle in the set of high-confidence obstacles based on the motion state using preset first screening conditions, so as to obtain the first obstacle set. The confidence analysis module is used to extract the pixel size information and position information of each low-confidence obstacle in the low-confidence obstacle set in the target frame image, and to calculate the size confidence and position confidence of each low-confidence obstacle based on the pixel size information and the position information. The second filtering module is used to determine the second obstacle in the low-confidence obstacle set based on the size confidence and position confidence using preset second filtering conditions, so as to obtain the second obstacle set; The collision index calculation module is used to merge the first obstacle set and the second obstacle set to generate a candidate obstacle target set, and to calculate the collision threat index between each obstacle target in the candidate obstacle target set and the intelligent driving vehicle. The obstacle determination module is used to determine the obstacle targets whose collision threat index is greater than a preset collision threat threshold as the target obstacle set.