Vehicle-mounted image recognition and target detection system based on deep learning

Through the collaborative optimization of distributed monitoring modules, multi-angle regional scene matching models, and decision-making modules, the problems of detection accuracy and trajectory control robustness of vehicle image recognition systems in unstructured environments have been solved, enabling safe obstacle avoidance and real-time control of autonomous driving in complex scenarios.

CN120953956BActive Publication Date: 2026-03-17BEIJING XINRUITE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing vehicle-mounted image recognition and target detection systems have high false detection and false negative rates in unstructured environments, making it difficult to implement safe avoidance strategies in unmarked or occluded environments. Furthermore, multimodal fusion algorithms lack robustness under low-quality input.

Method used

A distributed monitoring module is used for multimodal data fusion and cross-vehicle information sharing. Combined with a multi-angle regional scene matching model, a dynamically enhanced vehicle operation topology space is constructed. A label planning module is used for differentiated scene modeling and vehicle trajectory planning. The decision module uses particle swarm optimization algorithm and fuzzy control strategy for real-time trajectory control.

Benefits of technology

It significantly improves the target detection accuracy and trajectory control robustness of autonomous driving in occluded environments and unlabeled areas, reduces collision risks, and enables intelligent planning and real-time control in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953956B_ABST
    Figure CN120953956B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of vehicle control, and particularly relates to a vehicle-mounted image recognition and target detection system based on deep learning, comprising: a distributed monitoring module which collects target vehicle operation, obstacle and traffic signal information through multi-modal classification and scene matching, combines shared data to complete occlusion area marking and information pairing, and forms an enhanced monitoring set; a label planning module which constructs an enhanced topological space based on the enhanced monitoring set, adjusts the running trajectory under different scenes in combination with the vehicle and pedestrian trajectory probability distribution fed back by dynamic intention recognition; an action recognition module which predicts the trajectory parameters and collision probability of non-target vehicles and pedestrians by using Bayesian and multi-modal algorithms; and a decision module which generates real-time control instructions by particle swarm optimization and fuzzy control, and optimizes control parameters through simulation feedback; the application realizes intelligent trajectory planning and real-time control in complex scenes, and improves the detection accuracy and control robustness in occluded and unlabeled areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle control technology, and particularly relates to a deep learning-based vehicle image recognition and target detection system and system. Background Technology

[0002] With the rapid development of electric autonomous driving technology, in-vehicle image recognition and target detection systems have become core modules of environmental perception. Existing technologies, mostly based on deep learning methods, have made significant progress in structured road scenarios (such as highways and urban arterial roads), enabling safe navigation through traffic sign recognition, lane detection, and standardized obstacle prediction. However, in unstructured environments lacking traffic signs (such as school grounds, residential areas, and remote alleyways), traditional systems face severe challenges: First, blurred road boundaries and unclear traffic rules prevent vehicles from relying on preset signs for path decisions; second, the trajectories of dynamic targets such as pedestrians and non-motorized vehicles exhibit high randomness due to the lack of signal constraints, especially in sudden scenarios like "ghost pedestrians," where existing algorithms, relying on historical behavior patterns, struggle to quickly construct safe avoidance strategies; third, variable lighting and frequent occlusion in complex scenarios degrade the quality of visual sensor data, and multimodal fusion algorithms lack robustness under low-quality input. These problems significantly increase the false detection and false negative rates of existing systems in unstructured environments, hindering the full-domain application and safety improvement of autonomous driving technology. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention proposes a deep learning-based vehicle image recognition and target detection system, comprising: a distributed monitoring module that uses multimodal classification algorithms and multi-angle regional scene matching models to collect real-time target vehicle operation information, dynamic and static obstacles, and traffic signal label information, and combines shared information to complete occluded area classification labeling and information pairing, forming an enhanced monitoring set; a label planning module that constructs an enhanced topology space for the target area based on the enhanced monitoring set, and uses graph algorithms and trajectory planning models, combined with the vehicle and pedestrian intent probability distribution from the action recognition module, to dynamically adjust the running trajectory in scenarios with / without occlusion and with / without labels; an action recognition module that uses Bayesian algorithms and multimodal trajectory prediction algorithms to accurately predict the running trajectory parameters and collision probabilities of non-target vehicles and pedestrians; and a decision-making module that uses particle swarm optimization algorithms and fuzzy control algorithms to generate real-time optimal operation control commands and dynamically optimize the control parameter space through a simulation feedback mechanism. The system achieves intelligent planning and real-time control of vehicle running trajectories in complex scenarios through multi-algorithm fusion and multimodal data processing, significantly improving the target detection accuracy and trajectory control robustness of autonomous driving in occluded environments and unlabeled areas, and effectively reducing collision risks.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A deep learning-based vehicle image recognition and target detection system includes: a distributed monitoring module, a label planning module, an action recognition module, and a decision module;

[0006] The distributed monitoring module is used to collect target vehicle operation information, dynamic and static obstacle information, traffic signal tag information, and shared information of non-target vehicles in different directions within the same operating area in real time. Based on the collected information, it identifies and classifies the presence and absence of traffic signal tags and the state of obstruction. At the same time, it performs multi-angle shared information pairing of obstruction areas to obtain an enhanced monitoring set of target vehicles.

[0007] The label planning module is used to plan the vehicle's running trajectory in labeled or unlabeled areas under occlusion or unobstructed conditions based on the target vehicle enhanced monitoring set and the preset target vehicle trajectory planning model. At the same time, it uses the dynamic intent recognition model configured by the action recognition module to obtain the intended movement trajectory and movement intent probability distribution of corresponding vehicles and pedestrians in unobstructed and occluded areas, and feeds back the intended movement trajectory and movement intent probability distribution of vehicles and pedestrians to the target vehicle trajectory planning model to make real-time adjustments to the unobstructed and unobstructed autonomous vehicle running trajectory planning.

[0008] The decision module is used to input the adjusted trajectory planning information of the unmanned vehicles with and without obstructions into the operation control model to control and adjust the vehicle operation status in real time, so as to minimize the probability of vehicle collision.

[0009] Specifically, the distributed monitoring module includes monitoring units, sharing units, classification units, and pairing units;

[0010] The monitoring unit is used to collect target vehicle operation information, dynamic obstacle and static obstacle distribution information within the visible range, and traffic signal tag information in real time;

[0011] The sharing unit, based on a preset signal transmission node network, shares the acquired dynamic and static obstacle distribution information and traffic signal tag information within the visible range with target vehicles traveling in different directions within the same operating area.

[0012] The classification unit, based on the information collected by the monitoring unit and the information shared by the sharing unit, combined with a preset multimodal classification algorithm, obtains the feature space of collected data of different data types, and based on the feature space of collected data of different data types, uses an automatic labeling algorithm to label the unobstructed area, obstructed area and traffic signal label information corresponding to the target vehicle during the driving process.

[0013] The pairing unit, based on the occluded area information identified by the target vehicle and the collected shared information, obtains the multi-angle matching scene information and matching accuracy of the occluded area corresponding to the target vehicle through a multi-angle regional scene matching model.

[0014] Specifically, the label planning module includes topology modeling units and map layer units;

[0015] The topology modeling unit is used to construct an enhanced topology space for the target vehicle's operation based on the map layer information of the target vehicle's operating area, traffic signal label information, and the corresponding static and dynamic obstacle distribution information in the unobstructed and obstructed areas, through a graph algorithm.

[0016] The map layer unit obtains an enhanced topology space for the target vehicle's operation with location markers based on the target vehicle's operating topology space combined with map layer algorithms.

[0017] Specifically, the action recognition module includes a vehicle trajectory prediction unit and a pedestrian target trajectory unit;

[0018] The vehicle trajectory prediction unit is used to obtain the non-target vehicle running trajectory parameters corresponding to the different scene information based on the running information of the target vehicle within the visible range and the non-target vehicle in the shared information, respectively, by using a vehicle trajectory analysis and prediction algorithm combined with Bayesian algorithm.

[0019] The differential scene information includes the distribution information of dynamic and static obstacles, traffic signal label information, and operation information of non-target vehicles within the visible range corresponding to areas with traffic signal labels and no obstruction, areas with traffic signal labels and no obstruction, and areas with no traffic signal labels and obstruction.

[0020] The non-target vehicle trajectory parameters include all predicted trajectories, speeds, directions, and the probability of colliding with the target vehicle under the current trajectory, speed, and direction.

[0021] The pedestrian target trajectory unit is used to obtain the pedestrian trajectory parameters corresponding to the different scene information based on the pedestrian running trajectory and limb and gait movement information within the visible range of the target vehicle and in the shared information under the different scene information, through gait recognition algorithm combined with multimodal motion trajectory prediction algorithm.

[0022] The pedestrian trajectory parameters include the predicted trajectory and trajectory prediction accuracy at each moment, the direction and speed of travel, and the probability of collision with the target vehicle at each moment under the corresponding trajectory, direction and speed.

[0023] Specifically, the decision-making module includes a decision optimization unit, a decision control unit, and a simulation feedback unit;

[0024] The decision optimization unit is used to obtain the target vehicle's real-time optimal operating trajectory and control parameter space under different scene information based on the target vehicle's real-time operating information, non-target vehicle operating trajectory parameters under different scene information, pedestrian walking trajectory parameters, and static obstacle distribution information within the target vehicle's visible range, through a vehicle trajectory control algorithm optimized by particle swarm optimization.

[0025] The decision control unit generates real-time operation control commands for the target vehicle under different scenario information by combining the target vehicle's real-time optimal running trajectory and control parameter space with a fuzzy control algorithm.

[0026] The simulation feedback unit is used to simulate the control anomalies and collision probabilities of the target vehicle under the control of the generated real-time operation control commands, based on the enhanced topology space of the target vehicle's operation and the differential scene information, and to feed the simulation results back to the decision optimization unit to adjust the optimal operating trajectory and control parameter space of the target vehicle in real time, so as to minimize the vehicle collision probability.

[0027] Specifically, the construction and training process of the multi-angle regional scene matching model includes:

[0028] Extract the locally visible boundary image of the occluded area from the historical operation record of the target vehicle, as well as the LiDAR point cloud distribution information of the static environment and static visible markers within the locally visible boundary. Extract multi-angle panoramic images and shooting angles of the occluded area obtained by different non-target vehicles from the same timestamp of the shared information, as well as the LiDAR point cloud distribution information of the static environment and static visible markers within the locally visible boundary of the occluded area in the panoramic image.

[0029] Based on the static environment distribution and LiDAR point cloud distribution information of the target vehicle within the locally visible boundary and the corresponding static environment distribution and LiDAR point cloud distribution information of different non-target vehicles within the locally visible boundary at different perspectives in the panoramic image, feature correspondence is performed through timestamps, and matching information pairs are constructed through the aligned information.

[0030] Based on the matching information pairs, the Poisson surface reconstruction algorithm is used to generate a first mesh model of the locally visible boundary corresponding to the occluded area and a second mesh model of the locally visible boundary within the panoramic image corresponding to different viewpoints, and the corresponding first modeling loss and second modeling loss are obtained.

[0031] Based on the first and second mesh models, the Panoptic-DeepLab model is used to automatically segment and label the corresponding static markers within the locally visible boundaries, dynamic obstacle tags at the same time point, and traffic signal tags to obtain the first and second mesh models with complete labels.

[0032] Specifically, the construction and training process of the multi-angle region scene matching model also includes:

[0033] The first and second mesh models with complete annotations are input into the initial feature extraction layer of the multi-angle region scene matching model to obtain the first initial feature space and the second initial feature space.

[0034] The first initial feature space and the second initial feature space are synchronously and in parallel input to the point cloud feature extraction layer and the radar feature extraction layer, respectively, to obtain the first point cloud feature space, the second point cloud feature space, the first radar feature space, and the second radar feature space.

[0035] The first initial feature space and the second initial feature space, the first point cloud feature space and the second point cloud feature space, and the first radar feature space and the second radar feature space are input into a multi-level feature fusion layer combined with a Bayesian neural network to perform multi-level and multi-view feature fusion, thereby obtaining a multi-level and multi-view fused feature space and a loss function corresponding to each level of feature fusion.

[0036] The multi-level, multi-view fusion feature space and the loss function corresponding to each level of feature fusion, along with the first modeling loss and the second modeling loss, are input into the output layer combined with the cosine similarity function for training. This yields the matching accuracy of the local visible boundary of the occluded area of ​​the target vehicle and the corresponding multi-view panoramic image of the non-target vehicle at each time step, as well as the trained multi-angle region scene matching model.

[0037] Specifically, the process of constructing the target vehicle's enhanced topology space includes:

[0038] Based on the multi-angle regional scene matching model, obtain the matched multi-view panoramic image of the target vehicle at the same time point corresponding to the occluded area of ​​the target vehicle.

[0039] Based on the first and second initial feature spaces, the first and second point cloud feature spaces, the first and second radar feature spaces, and the dynamic and static obstacle information and traffic signal tag information within the target vehicle's visible range, an enhanced 3D space for the target vehicle is obtained by combining a map-layer 3D modeling algorithm.

[0040] Based on the dynamic and static obstacle information and traffic signal location information of each location point within the visible range of the target vehicle at the current time point in the enhanced three-dimensional space of the target vehicle, and the dynamic obstacle information and traffic signal location information of each location point in the occluded area, the obstacle node and signal node in the enhanced node space corresponding to the target vehicle at each time point are constructed.

[0041] Simultaneously, a first connection relationship is constructed based on the distance between the target vehicle and the corresponding static obstacle within the visible range at the same time point, and a second connection relationship is constructed based on the product of the distance between the target vehicle and the corresponding dynamic obstacle within the visible range at the same time point, the probability of collision, and the probability of the presence of a traffic signal tag within the visible range.

[0042] Specifically, the process of constructing the target vehicle's enhanced topology space also includes:

[0043] The third connection relationship is constructed by multiplying the distance between the target vehicle and the corresponding vehicle in the occluded area at the same time point, the probability of collision, the probability of traffic signal in the occluded area, and the matching accuracy of the occluded area at the current time.

[0044] Simultaneously, a fourth connection relationship is constructed by multiplying the distance between the target vehicle and the pedestrian in the corresponding occluded area at the same time point, the accuracy of the trajectory prediction, the probability of a collision, the probability of a traffic signal in the occluded area, and the matching accuracy of the occluded area at the current time.

[0045] Based on the first, second, third, and fourth connection relationships of obstacle nodes and signal nodes in the augmented node space corresponding to the target vehicle at each time point, a static target vehicle augmented topology space corresponding to each time point is constructed by a graph algorithm.

[0046] Specifically, the process of constructing the target vehicle's enhanced topology space also includes:

[0047] Based on the static target vehicle enhanced topology space, vehicle trajectory analysis and prediction algorithms are configured on the second and third connection relationships, and gait recognition algorithms are configured on the fourth connection relationship. Combined with multimodal motion trajectory prediction algorithms and three-dimensional simulation algorithms, the dynamic operation simulation of the target vehicle in areas with traffic signals, without traffic signals, and with and without obstruction is carried out at continuous time points to obtain the target vehicle operation enhanced topology space.

[0048] Compared with the prior art, the beneficial effects of the present invention are:

[0049] This invention addresses the shortcomings of existing technologies by employing a multimodal data fusion and cross-vehicle information sharing mechanism within a distributed monitoring module. Combined with a multi-angle regional scene matching model, it performs multi-source perception compensation for occluded areas, constructing a dynamically enhanced vehicle operation topology space and effectively improving global perception capabilities in unstructured environments. Based on the collaborative optimization of differentiated scene modeling and vehicle trajectory planning models within the label planning module, it generates feasible trajectories that conform to the physical constraints of complex scenarios by real-time fusion of Bayesian trajectory prediction of dynamic obstacles and multimodal pedestrian intent recognition results. The decision module uses dynamic coupling of particle swarm optimization algorithm and fuzzy control strategy to achieve online adaptive adjustment of trajectory control parameters. Combined with a simulation feedback mechanism, it performs forward-looking assessment and closed-loop correction of potential risks, forming a full-link collaborative optimization of perception, planning, and decision-making. This significantly enhances the system's real-time obstacle avoidance robustness and trajectory smoothness in unmarked and multi-occluded scenarios. Attached Figure Description

[0050] Figure 1 This is a flowchart of the deep learning-based vehicle image recognition and target detection system of the present invention.

[0051] Figure 2 This is a schematic diagram of the occlusion area of ​​the present invention. Detailed Implementation

[0052] Please see Figure 1 The present invention provides an embodiment of a deep learning-based vehicle image recognition and target detection system, comprising the following steps: a distributed monitoring module, a label planning module, an action recognition module, and a decision module;

[0053] The distributed monitoring module is used to collect target vehicle operation information, dynamic and static obstacle information, traffic signal tag information, and shared information of non-target vehicles in different directions within the same operating area in real time. Based on the collected information, it identifies and classifies the presence and absence of traffic signal tags and the state of obstruction. At the same time, it performs multi-angle shared information pairing of obstruction areas to obtain an enhanced monitoring set of target vehicles.

[0054] The label planning module is used to plan the vehicle's running trajectory in labeled or unlabeled areas under occlusion or unobstructed conditions based on the target vehicle enhanced monitoring set and the preset target vehicle trajectory planning model. At the same time, it uses the dynamic intent recognition model configured by the action recognition module to obtain the intended movement trajectory and movement intent probability distribution of corresponding vehicles and pedestrians in unobstructed and occluded areas, and feeds back the intended movement trajectory and movement intent probability distribution of vehicles and pedestrians to the target vehicle trajectory planning model to make real-time adjustments to the unobstructed and unobstructed autonomous vehicle running trajectory planning.

[0055] The decision module is used to input the adjusted trajectory planning information of the unmanned vehicles with and without obstructions into the operation control model to control and adjust the vehicle operation status in real time, so as to minimize the probability of vehicle collision.

[0056] Furthermore, the distributed monitoring module in this embodiment includes a monitoring unit, a sharing unit, a classification unit, and a pairing unit;

[0057] The monitoring unit is used to collect target vehicle operation information, dynamic obstacle and static obstacle distribution information within the visible range, and traffic signal tag information in real time;

[0058] Furthermore, the target vehicle operation information in this embodiment comprehensively covers positioning attitude, dynamic state, control commands, and environmental perception data: positioning and attitude are obtained through high-precision GPS to acquire latitude, longitude, and altitude, and combined with data from wheel speed sensors and inertial measurement units to analyze tire slip ratio and three-axis angular velocity; dynamic state includes longitudinal velocity and acceleration, lateral offset, and yaw rate, reflecting the vehicle's motion characteristics; control command feedback integrates steering wheel angle, pedal opening, and gear status to form a closed loop of driving operation. Environmental perception data is divided into dynamic obstacle and static obstacle analysis. Dynamic obstacle information includes the position trajectory prediction and attribute recognition of non-target vehicles (type, turn signal, and license plate), skeletal key points and intent classification of pedestrians and non-motorized vehicles, and multimodal detection of special targets (animals, drones); static obstacles include the size and traversability score of fixed road objects, terrain slope and abnormal road surface area recognition (potholes, water accumulation), target detection in construction areas, and temporary obstacle reconstruction; traffic signal labels analyze traffic light status (color, countdown), road sign semantics (speed limit, prohibition), and rule inference of unmarked areas (probability of right-hand passage, pedestrian priority area). Global information is used to accurately describe the vehicle's motion state and perceive all elements of complex environments through multi-sensor fusion and deep learning models, supporting intelligent decision-making and path planning.

[0059] The sharing unit, based on a preset signal transmission node network, shares the acquired dynamic and static obstacle distribution information and traffic signal tag information within the visible range with target vehicles traveling in different directions within the same operating area.

[0060] Furthermore, the technical implementation process of the corresponding shared unit in this embodiment includes:

[0061] First, dynamic obstacles are structured and encoded using the Protobuf protocol, and efficient transmission is achieved by utilizing lightweight serialization and semantic tag classification.

[0062] Second, static obstacles are compressed using the Octree spatial segmentation algorithm, and their passability is quantified by combining reflection intensity classification and decision tree model.

[0063] Third, the traffic signal uses state machine coding to map discrete event logic to ensure real-time response;

[0064] Fourth, based on the IEEE 1588 precision clock synchronization protocol, timing deviations of multi-source data are eliminated, and SLAM real-time pose data drives UTM coordinate system projection to achieve spatial consistency alignment.

[0065] Fifth, the DSRC and 5GNR hybrid network uses reinforcement learning for dynamic routing optimization, hierarchical coding strategy to ensure low latency for emergency data and high reliability for regular data transmission, and Merkle tree verification chain and blockchain notarization to ensure data integrity and traceability.

[0066] Sixth, the receiving end fuses multi-source sensing results based on DS evidence theory and dynamically weights them to improve confidence; incremental Bayesian filtering updates the occupied grid map and embeds shared obstacle information;

[0067] Seventh, the spatiotemporal causal reasoning engine constructs an obstacle evolution map, eliminates false associations through Bayesian networks, and deduces behavioral intentions.

[0068] The classification unit, based on the information collected by the monitoring unit and the information shared by the sharing unit, combined with a preset multimodal classification algorithm, obtains the feature space of collected data of different data types, and based on the feature space of collected data of different data types, uses an automatic labeling algorithm to label the unobstructed area, obstructed area and traffic signal label information corresponding to the target vehicle during the driving process.

[0069] It should be noted that the classification unit in this embodiment corresponds to a more detailed implementation process, including:

[0070] Based on the multimodal sensor information collected by the monitoring unit and the information shared by the sharing unit, time synchronization and coordinate alignment are performed on LiDAR point cloud, visual image and radar data to construct a spatiotemporal reference.

[0071] Geometric features are extracted through LiDAR point cloud clustering, semantic information is obtained through visual target detection, and motion attributes are analyzed using radar Doppler features. This constructs feature spaces for different data types, specifically including:

[0072] 1) Kalman filtering is used to fuse centroid and velocity data from multiple sensors, and the ground point cloud is segmented using the Region Growing algorithm and the material is classified by reflection intensity;

[0073] 2) Simultaneously, the SSD model is used to analyze traffic signal colors and road signs, and V2X collaborative verification is used to improve recognition robustness;

[0074] 3) Based on point cloud density, visual consistency and multi-sensor speed verification, determine whether there is an occluded area. Then, use ray casting and multi-view Poisson reconstruction to complete the 3D structure of the occluded area. Finally, use a temporal hidden Markov model filter to eliminate transient noise between the unoccluded area and the occluded area.

[0075] 4) Derive access rules for unlabeled areas using historical trajectory clustering and graph neural networks;

[0076] Third, based on the feature space of the collected data, the unobstructed area, obstructed area, and traffic signal label information corresponding to the target vehicle during the driving process are labeled by an automatic labeling algorithm. Embedded real-time reasoning and automatic labeling of different areas and traffic signals are realized through a knowledge distillation compression model.

[0077] Furthermore, for a more detailed explanation of the occluded and unoccluded areas in this embodiment, please refer to [link to relevant documentation]. Figure 2 Let G and H represent the normal operating directions of two intersecting roads. Let A be the target vehicle, F be a stationary truck within A's field of vision, and the dashed box corresponding to D be a roadside building that blocks the target vehicle A's view of the road segment corresponding to the non-target vehicle B. Pedestrian C is the pedestrian corresponding to the road segment corresponding to the non-target vehicle B, which is a road segment with traffic signals. The road segment corresponding to the non-target vehicle I is an alley without traffic signals, which is also an obstruction area for the target vehicle A. E is the pedestrian corresponding to the road segment of the target vehicle A, but its path is blocked by the non-target vehicle F. Therefore, the road segment corresponding to the non-target vehicle B and the part of the path corresponding to pedestrian E are obstruction areas. The area where pedestrian E is obstructed by the non-target vehicle F is defined here as the high-incidence obstruction area of ​​"ghost peek".

[0078] The pairing unit, based on the occluded area information identified by the target vehicle and the collected shared information, obtains the multi-angle matching scene information and matching accuracy of the occluded area corresponding to the target vehicle through a multi-angle regional scene matching model.

[0079] Furthermore, the construction and training process of the multi-angle region scene matching model in this embodiment includes:

[0080] Extract the locally visible boundary image of the occluded area from the historical operation record of the target vehicle, as well as the LiDAR point cloud distribution information of the static environment and static visible markers within the locally visible boundary. Extract multi-angle panoramic images and shooting angles of the occluded area obtained by different non-target vehicles from the same timestamp of the shared information, as well as the LiDAR point cloud distribution information of the static environment and static visible markers within the locally visible boundary of the occluded area in the panoramic image.

[0081] Based on the static environment distribution and LiDAR point cloud distribution information of the target vehicle within the locally visible boundary and the corresponding static environment distribution and LiDAR point cloud distribution information of different non-target vehicles within the locally visible boundary at different perspectives in the panoramic image, feature correspondence is performed through timestamps, and matching information pairs are constructed through the aligned information.

[0082] Furthermore, in this embodiment, the target vehicle acquires locally visible boundary images, static environment point clouds, and multi-angle panoramic data of neighboring vehicles through multi-modal sensors, and aligns the timestamps of multi-source data using a precision clock synchronization protocol; based on real-time dynamic positioning technology and inertial measurement unit data, the data is unified to the vehicle coordinate system through geographic projection and rigid body transformation; the static environment point cloud is input into the panoramic segmentation model for semantic annotation, and material attributes are classified in combination with reflection intensity; the dynamic obstacle point cloud undergoes iterative nearest-point registration to verify geometric consistency, and semantic tags are fused to resolve logical conflicts; finally, the multi-view observation data of the vehicle and neighboring vehicles are integrated, and a holographic model of the occluded area is generated through a 3D reconstruction algorithm, forming a technical closed loop of "multi-source acquisition, spatiotemporal alignment, semantic segmentation, dynamic verification, and 3D modeling", achieving centimeter-level accuracy in environmental perception and blind spot completion.

[0083] Based on the matching information pairs, the Poisson surface reconstruction algorithm is used to generate a first mesh model of the locally visible boundary corresponding to the occluded area and a second mesh model of the locally visible boundary within the panoramic image corresponding to different viewpoints, and the corresponding first modeling loss and second modeling loss are obtained.

[0084] Furthermore, the construction process of the first Mesh model and the second Mesh model in this embodiment is as follows:

[0085] By filtering and noise reduction, ground segmentation, semantic annotation and multi-view registration, the output point cloud data with registration alignment and complete semantic annotation is provided to provide spatially consistent multi-source input for subsequent reconstruction.

[0086] Based on the normal constraints and spatial fusion of preprocessed point clouds, single-view fine mesh and multi-view fused mesh of the target vehicle are generated respectively.

[0087] Through geometric, semantic, and multi-view consistency evaluation, the parameters of the Mesh model are dynamically corrected and local reconstruction is triggered, ultimately outputting the optimized first Mesh model and second Mesh model.

[0088] Based on the first and second mesh models, the Panoptic-DeepLab model is used to automatically segment and label the corresponding static markers within the locally visible boundaries, dynamic obstacle tags at the same time point, and traffic signal tags to obtain the first and second mesh models with complete labels.

[0089] The first and second mesh models with complete annotations are input into the initial feature extraction layer of the multi-angle region scene matching model to obtain the first initial feature space and the second initial feature space.

[0090] Furthermore, the process of obtaining the first initial feature space and the second initial feature space in this embodiment includes:

[0091] The first and second mesh models, which label static markers, dynamic obstacles, and traffic signal tags, are divided into regular three-dimensional voxel grids to obtain voxelized grid data, where each voxel contains geometric and semantic information.

[0092] Furthermore, in this embodiment, each voxel stores geometric attributes and semantic labels, and the continuous surface model is converted into structured mesh data through spatial discretization.

[0093] The point cloud distribution density within each voxel is statistically analyzed and normalized. The covariance matrix of the local point cloud is analyzed to calculate the surface curvature. The normal vector direction is extracted by principal component analysis to quantify the local geometric characteristics and obtain the geometric feature vector of each voxel, such as density, curvature, and normal vector.

[0094] Based on voxel data and visible light images, semantic segmentation and label mapping are used to output semantic feature vectors containing category information;

[0095] Based on geometric and semantic feature vectors, the first and second initial feature spaces are output through concatenation, context encoding, and coordinate alignment.

[0096] The first initial feature space and the second initial feature space are synchronously and in parallel input to the point cloud feature extraction layer and the radar feature extraction layer, respectively, to obtain the first point cloud feature space, the second point cloud feature space, the first radar feature space, and the second radar feature space.

[0097] Furthermore, the process of obtaining the first point cloud feature space and the second point cloud feature space in this embodiment includes:

[0098] Based on the original point cloud, key points are selected by sampling the farthest point and neighborhood geometric features are extracted, and the key points and their local feature vectors are output.

[0099] Based on local feature vectors, a wider range of structural features is extracted by expanding the neighborhood range and enhancing the multilayer perceptron, and a mid-layer feature vector is output.

[0100] Based on the mid-level feature vector, global pooling and multi-layer perceptron are used to capture the overall geometric characteristics, output global features, and form a complete representation of the point cloud.

[0101] Based on local feature vectors, mid-level feature vectors and global features, redundant information is compressed by concatenation and fully connected layers, and a high-dimensional vector that integrates multi-scale features is output.

[0102] Based on high-dimensional vectors of target, shared viewpoint features and multi-scale features, spatial alignment is achieved through coordinate transformation to obtain the first point cloud feature space and the second point cloud feature space.

[0103] Furthermore, the process of obtaining the first radar feature space and the second radar feature space in this embodiment includes:

[0104] The original four-dimensional radar data is flattened into a one-dimensional sequence. By fusing the four-dimensional radar data, a four-dimensional radar data point space is obtained. It should be further noted that each data point in this embodiment includes, but is not limited to, the original features of four dimensions: range, azimuth, elevation, and Doppler velocity.

[0105] For each data point's three-dimensional spatial coordinates, a 512-dimensional position embedding is generated using a sine function. Specifically, sine and cosine values ​​of different frequencies are independently calculated for each coordinate axis and alternately concatenated to form a position encoding vector. During the encoding process, the wavelength parameter is set to 100 meters to ensure that the encoding range covers the typical detection range of vehicle-mounted radar.

[0106] Using eight parallel attention heads, each attention head embeds a 512-dimensional location into a 64-dimensional query, key, and value vector. Attention weights are calculated through dot product, and global contextual information is fused.

[0107] Global average pooling is performed on the fused global context information output by the Transformer to extract feature statistics for all spatial locations, generating a 512-dimensional global feature vector.

[0108] By using a fully connected layer, the 512-dimensional global feature vector is linearly mapped to 256 dimensions, reducing the feature dimension while retaining key information, and obtaining the first radar feature space and the second radar feature space.

[0109] It should be further explained that, in this embodiment, the first radar feature space represents the motion pattern of the dynamic target from the perspective of the target vehicle; the second radar feature space represents the dynamic target information detected by encoding the shared perspective, and is spatiotemporally aligned with the first feature space.

[0110] The first initial feature space and the second initial feature space, the first point cloud feature space and the second point cloud feature space, and the first radar feature space and the second radar feature space are input into a multi-level feature fusion layer combined with a Bayesian neural network to perform multi-level and multi-view feature fusion, thereby obtaining a multi-level and multi-view fused feature space and a loss function corresponding to each level of feature fusion.

[0111] The multi-level feature fusion layer includes a multi-view cross-modal feature fusion sub-layer, a multi-view geometric fusion sub-layer, and a multi-view semantic fusion sub-layer. Furthermore, in this embodiment, the multi-view semantic fusion sub-layer is preferably Multilingual-BERT.

[0112] Furthermore, the implementation process of the multi-view cross-modal feature fusion sublayer in this embodiment includes:

[0113] First, the three types of heterogeneous features—initial feature space, point cloud feature space, and radar feature space—are defined as graph nodes. The initial feature space contains the basic features of the target vehicle and the shared viewpoint, the point cloud feature space covers three-dimensional geometric structure information, and the radar feature space represents the dynamic target motion mode. This results in the output of a basic graph structure containing multimodal nodes.

[0114] Second, edge connection rules are established based on the spatial proximity (such as node association within the Euclidean distance threshold) and temporal consistency (continuous frame feature change threshold) between multimodal graph nodes, thereby enhancing the semantic relationship of the graph structure through spatiotemporal association;

[0115] Third, the node features are projected and dimensionality reduced based on the learnable weight matrix, and the attention coefficients between nodes are calculated based on the projected features to quantify the interaction strength of different modal features and output low-dimensional projected features and dynamic relationship weights.

[0116] Fourth, a multi-head attention network is used to process the projected features and dynamic relationship weights. Complex associations are captured through a multi-branch interaction mode. The results of each branch are concatenated to form a cross-modal joint feature vector that integrates multiple interaction modes. The multi-head attention network is preferably the multi-head attention mechanism in the transformer algorithm.

[0117] Fifth, based on the joint feature vector, the similarity of matching nodes of the same type is maximized by comparing the loss function, while the temperature coefficient is dynamically adjusted to optimize the feature distribution. Finally, the optimized cross-modal fusion feature space is output, realizing the deep fusion and semantic enhancement of heterogeneous features under multiple perspectives.

[0118] Furthermore, the implementation process of the multi-view geometric fusion sublayer and the multi-view semantic fusion sublayer in this embodiment includes:

[0119] The input point cloud data is preprocessed, and key points in high curvature regions are selected by curvature calculation. At the same time, voxel grid filtering is used to suppress redundant points, thereby obtaining a representative set of key points that can characterize the geometric features of the point cloud.

[0120] The geometric histogram descriptor is calculated for the set. By statistically analyzing and normalizing the spatial distribution characteristics of the neighborhood of key points, a standardized feature descriptor vector is formed. Then, based on the descriptor vector, the Random Sample Consensus Algorithm (RANSAC) is applied. Through multiple random sampling iterations, the optimal initial transformation matrix is ​​calculated and selected to obtain the coarse registration result and the corresponding set of interior points.

[0121] Third, based on the coarse registration results, a robust kernel function is introduced to optimize the reprojection error, and the weight parameters are dynamically adjusted to reduce the influence of outliers, thereby outputting the optimized transformation matrix.

[0122] Fourth, based on the optimized transformation matrix, the position error and orientation error are weighted and summed according to preset weights to obtain a quantified registration quality evaluation index.

[0123] Fifth, for multi-view node features, the neighborhood aggregation operation of graph neural networks is used to fuse the features of adjacent nodes, and the feature dimensions are simplified through a hierarchical compression algorithm. At the same time, semantic label information is embedded to generate semantically enhanced graph node features.

[0124] Sixth, random weight sampling is performed on the semantically enhanced graph node features, and combined with the evidence lower bound (ELBO) loss constraint, the semantic probability distribution with uncertainty is derived to obtain the confidence of the quantitative classification result;

[0125] Seventh, based on semantic probability distribution, a cross-entropy loss function is constructed and an uncertainty weighting mechanism is introduced to set a preset penalty for low-confidence predictions, thereby outputting a robust semantic fusion result and realizing the joint optimization of geometric registration and semantic enhancement of multi-view point cloud data.

[0126] The multi-level, multi-view fusion feature space and the loss function corresponding to each level of feature fusion, along with the first modeling loss and the second modeling loss, are input into the output layer combined with the cosine similarity function for training. This yields the matching accuracy of the local visible boundary of the occluded area of ​​the target vehicle and the corresponding multi-view panoramic image of the non-target vehicle at each time step, as well as the trained multi-angle region scene matching model.

[0127] Furthermore, the label planning module in this embodiment includes a topology modeling unit and a map layer unit;

[0128] The topology modeling unit is used to construct an enhanced topology space for the target vehicle's operation based on the map layer information of the target vehicle's operating area, traffic signal label information, and the static and dynamic obstacle information corresponding to the unobstructed and obstructed areas, through a graph algorithm.

[0129] The map layer unit obtains an enhanced topology space for the target vehicle's operation with location markers based on the target vehicle's operating topology space combined with map layer algorithms.

[0130] Furthermore, the action recognition module in this embodiment includes a vehicle trajectory prediction unit and a pedestrian target trajectory unit;

[0131] The vehicle trajectory prediction unit is used to obtain the non-target vehicle running trajectory parameters corresponding to the different scene information based on the running information of the target vehicle within the visible range and the non-target vehicle in the shared information, respectively, by using a vehicle trajectory analysis and prediction algorithm combined with Bayesian algorithm.

[0132] It should be further explained that the differential scene information in this embodiment includes the distribution information of dynamic and static obstacles, traffic signal label information, and non-target vehicle operation information within the visible range corresponding to areas with traffic signal labels and no obstruction, areas with traffic signal labels and no obstruction, and areas with no traffic signal labels and obstruction.

[0133] The non-target vehicle trajectory parameters include all predicted trajectories, speeds, directions, and the probability of colliding with the target vehicle under the current trajectory, speed, and direction.

[0134] The pedestrian target trajectory unit is used to obtain pedestrian trajectory parameters corresponding to the different scene information based on the pedestrian running trajectory and limb and gait movement information within the visible range of the target vehicle and in the shared information under the different scene information, through gait recognition algorithm combined with multimodal motion trajectory prediction algorithm.

[0135] Furthermore, the workflow of the pedestrian target trajectory unit in this embodiment includes:

[0136] Based on signal status and occlusion conditions, the scene is classified into four categories. The characteristics of traffic signal effectiveness, occlusion area determination and obstacle distribution are extracted to obtain the differentiated scenes and scene feature vectors.

[0137] Furthermore, the different scenarios in this embodiment include: Scenario A is an area with traffic signal signs and no obstruction, Scenario B is an area with traffic signal signs and obstruction, Scenario C is an area without traffic signal signs and no obstruction, and Scenario D is an area without traffic signal signs and obstruction.

[0138] Based on scene classification labels and scene feature vectors, pedestrian posture is extracted through 2D / 3D key point detection. Combined with temporal filtering and smoothing data, gait cycle, stride length, and torso tilt angle parameters are quantified to obtain pedestrian gait features and intent judgment results.

[0139] Based on the fusion of basic motion parameters and scene extended features by scene features and gait features, the model architecture is adjusted according to scene type to obtain the probability distribution and confidence of pedestrian future trajectory.

[0140] Based on the pedestrian trajectory prediction results, physical constraints are applied according to the scene type, combined with dynamic risk scoring and non-target vehicle information to optimize the trajectory, thereby obtaining the corresponding pedestrian trajectory parameters under different scene information.

[0141] Furthermore, the decision-making module in this embodiment includes a decision optimization unit, a decision control unit, and a simulation feedback unit;

[0142] The decision optimization unit is used to obtain the target vehicle's real-time optimal operating trajectory and control parameter space under different scene information based on the target vehicle's real-time operating information, non-target vehicle operating trajectory parameters under different scene information, pedestrian walking trajectory parameters, and static obstacle distribution information within the target vehicle's visible range, through a vehicle trajectory control algorithm optimized by particle swarm optimization.

[0143] The decision control unit generates real-time operation control commands for the target vehicle under different scenario information by combining the target vehicle's real-time optimal running trajectory and control parameter space with a fuzzy control algorithm.

[0144] The simulation feedback unit is used to simulate the control anomalies and collision probabilities of the target vehicle under the control of the generated real-time operation control commands, based on the enhanced topology space of the target vehicle's operation and the differential scene information, and to feed the simulation results back to the decision optimization unit to adjust the optimal operating trajectory and control parameter space of the target vehicle in real time, so as to minimize the vehicle collision probability.

[0145] Furthermore, the process of constructing the target vehicle's enhanced topology space in this embodiment includes:

[0146] Based on the multi-angle regional scene matching model, obtain the matched multi-view panoramic image of the target vehicle at the same time point corresponding to the occluded area of ​​the target vehicle.

[0147] Based on the first and second initial feature spaces, the first and second point cloud feature spaces, the first and second radar feature spaces, and the dynamic and static obstacle information and traffic signal tag information within the target vehicle's visible range, an enhanced 3D space for the target vehicle is obtained by combining a map-layer 3D modeling algorithm.

[0148] Based on the dynamic and static obstacle information and traffic signal location information of each location point within the visible range of the target vehicle at the current time point in the enhanced three-dimensional space of the target vehicle, and the dynamic obstacle information and traffic signal location information of each location point in the occluded area, the obstacle node and signal node in the enhanced node space corresponding to the target vehicle at each time point are constructed.

[0149] Simultaneously, a first connection relationship is constructed based on the distance between the target vehicle and the corresponding static obstacle within the visible range at the same time point, and a second connection relationship is constructed based on the product of the distance between the target vehicle and the corresponding dynamic obstacle within the visible range at the same time point, the probability of collision, and the probability of the presence of a traffic signal tag within the visible range.

[0150] The third connection relationship is constructed by multiplying the distance between the target vehicle and the corresponding vehicle in the occluded area at the same time point, the probability of collision, the probability of traffic signal in the occluded area, and the matching accuracy of the occluded area at the current time.

[0151] Simultaneously, a fourth connection relationship is constructed by multiplying the distance between the target vehicle and the pedestrian in the corresponding occluded area at the same time point, the accuracy of the trajectory prediction, the probability of a collision, the probability of a traffic signal in the occluded area, and the matching accuracy of the occluded area at the current time.

[0152] Based on the first, second, third, and fourth connection relationships of obstacle nodes and signal nodes in the augmented node space corresponding to the target vehicle at each time point, a static target vehicle augmented topology space corresponding to each time point is constructed by a graph algorithm.

[0153] Based on the static target vehicle enhanced topology space, vehicle trajectory analysis and prediction algorithms are configured on the second and third connection relationships, and gait recognition algorithms are configured on the fourth connection relationship. Combined with multimodal motion trajectory prediction algorithms and three-dimensional simulation algorithms, the dynamic operation simulation of the target vehicle in areas with traffic signals, without traffic signals, and with and without obstruction is carried out at continuous time points to obtain the target vehicle operation enhanced topology space.

[0154] It should be further explained that the process of constructing the target vehicle's enhanced topology space in this embodiment includes:

[0155] Based on vehicle multi-source sensor data, a multi-angle regional scene matching model is used to fuse multi-view information of the occluded area and output a panoramic image of the occluded area of ​​the target vehicle.

[0156] Based on the panoramic image, the initial feature space, point cloud feature space and radar feature space are extracted, and mapped to the target vehicle coordinate system through coordinate transformation, outputting a multimodal feature space in a unified coordinate system;

[0157] Based on a multimodal feature space with a unified coordinate system and a high-precision map, the system aligns the outlines of static obstacles with the positions of dynamic targets through rasterization, and outputs enhanced data that integrates the map with the real-time scene.

[0158] Based on the augmented data of the fused map and real-time scene, it is converted into a three-dimensional voxel mesh and stored geometric, semantic and dynamic attributes, and output a structured voxel space model.

[0159] Based on the voxel space model, corresponding nodes are generated and attributes are stored for static obstacles, dynamic targets and traffic lights, and the output is a set of nodes containing position, type and dynamic parameters;

[0160] Based on the node set, the four types of connection weights are dynamically calculated by combining distance, collision probability, signal influence and matching confidence, and a weighted node interaction relationship network is output.

[0161] Based on node sets and connection relationships, a graph structure is constructed and edge weights are optimized to output a spatiotemporally consistent static target vehicle enhanced topology space.

[0162] Based on the static target vehicle enhanced topology space and real-time sensor data, the node attributes are updated through trajectory prediction, occlusion completion and game theory model to output a dynamically evolving topology space;

[0163] Based on dynamic topology space, a path search algorithm is used to avoid high-risk edges and combined with signal status to simulate driving strategies, outputting safe paths and decision instructions;

[0164] Based on path and topology data, the system displays the distribution of node and edge weights through 3D rendering, marks high-risk areas in real time, and outputs a visual driving assistance interface.

[0165] This embodiment constructs a high-precision perception benchmark through hardware-level spatiotemporal synchronization and multi-source data fusion. It generates joint feature vectors using geometric, semantic, and motion feature extraction models and distinguishes between static and dynamic occlusion areas using semi-supervised learning. It compensates for perception blind spots based on multi-view reconstruction and uncertainty modeling, and dynamically adjusts risk levels through probabilistic virtual nodes. At the planning layer, it deeply integrates traffic signal status, occlusion conditions, and pedestrian biometrics, and uses a temporal attention mechanism and interactive topology modeling to achieve physically interpretable trajectory prediction. At the decision layer, it optimizes path strategies by dynamically enhancing the topological spatial encoding environment interaction relationships and combining risk quantification algorithms and closed-loop verification mechanisms. The end-to-end design introduces data integrity verification, adaptive fusion weight allocation, and online model compression technology, and combines a continuous learning framework to achieve adaptive capabilities to new scenarios and unknown obstacles, forming a closed-loop enhancement system from multimodal perception to game-theoretic decision-making.

[0166] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments under the guidance of the present invention without departing from the spirit and scope of the claims. All of these variations are within the protection scope of the present invention.

Claims

1. A vehicle-mounted image recognition and target detection system based on deep learning, characterized in that, The application relates to a distributed monitoring system for unmanned vehicles. The distributed monitoring module is used for collecting target vehicle running information, dynamic and static obstacle information, traffic signal label information and shared information of non-target vehicles in different running directions in the same running area in real time, and identifying and classifying the traffic signal label and the shielding state according to the collected information, and meanwhile, the shielding area multi-angle shared information is matched to obtain a target vehicle enhanced monitoring set. The label planning module is used for planning a vehicle running track in a labeled or unlabeled area with shielding or without shielding according to the target vehicle enhanced monitoring set and a preset target vehicle track planning model, and simultaneously, a dynamic intention recognition model configured by the action recognition module is used to obtain an intention running track and an intention probability distribution of a corresponding vehicle and pedestrian in the unlabeled area and the labeled area, and the intention running track and the intention probability distribution are fed back to the target vehicle track planning model to adjust the running track planning of the unmanned vehicle with shielding and without shielding in real time. The decision module is used for inputting the adjusted running track planning information of the unmanned vehicle with shielding and without shielding into a running control model to control and adjust the vehicle running state in real time so that the vehicle collision probability is minimized. The distributed monitoring module comprises a monitoring unit, a sharing unit, a classification unit and a matching unit. 2.The deep learning based vehicle-mounted image recognition and target detection system according to claim 1, wherein, The monitoring unit is used for collecting target vehicle running information, dynamic obstacle distribution information and static obstacle distribution information in a visual range and traffic signal label information in real time. The sharing unit shares the dynamic obstacle distribution information, the static obstacle distribution information and the traffic signal label information in the visual range to target vehicles in different running directions in the same running area based on a preset signal transmission node network. The classification unit obtains a collection data feature space of different data types based on the information collected by the monitoring unit and the information shared by the sharing unit and a preset multi-modal classification algorithm, and labels the unlabeled area, the labeled area and the traffic signal label information corresponding to the target vehicle in the driving process through an automatic labeling algorithm based on the collection data feature space of different data types. The matching unit obtains multi-angle matching scene information and a matching accuracy rate of the unlabeled area corresponding to the target vehicle through a multi-angle area scene matching model based on the shielding area information recognized by the target vehicle and the collected shared information. The label planning module comprises a topological modeling unit and a map layer unit. 3.The deep learning based vehicle-mounted image recognition and target detection system of claim 2, wherein, The topological modeling unit is used for constructing a target vehicle running topological space through a graph algorithm based on target vehicle running area map layer information, traffic signal label information, static obstacle distribution information and dynamic obstacle information in the unlabeled area and the labeled area. The map layer unit obtains a target vehicle running enhanced topological space with position marks based on the target vehicle running topological space and a map layer algorithm. The action recognition module comprises a vehicle track prediction unit and a pedestrian target track unit. 4.The deep learning based vehicle-mounted image recognition and target detection system of claim 3, wherein, ​ The vehicle trajectory prediction unit is configured to obtain non-target vehicle running trajectory parameters corresponding to the difference scenario information by combining a vehicle trajectory analysis prediction algorithm based on the Bayesian algorithm, according to the non-target vehicle running information in the shared information within the visual range of the target vehicle corresponding to the difference scenario information. The difference scenario information includes dynamic obstacle and static obstacle distribution information, traffic signal label information and non-target vehicle running information in the visual range corresponding to an unobstructed area with a traffic signal label, an obstructed area with a traffic signal label, an unobstructed area without a traffic signal label and an obstructed area without a traffic signal label. The non-target vehicle running trajectory parameters include all predicted running trajectories of the non-target vehicle, running speed, direction and probability of collision with the target vehicle under the current running trajectory and speed and direction. The pedestrian target trajectory unit is configured to obtain pedestrian running trajectory parameters corresponding to the difference scenario information by combining a multi-modal motion trajectory prediction algorithm based on a gait recognition algorithm, according to the pedestrian running trajectory and limb and gait motion information in the shared information within the visual range of the target vehicle corresponding to the difference scenario information. The pedestrian running trajectory parameters include a predicted running trajectory at each time, a running direction and speed, and a probability of collision with the target vehicle under the corresponding running trajectory, running direction and speed at each time. 5.The deep learning based vehicle-mounted image recognition and target detection system of claim 4, wherein, The decision module includes a decision optimization unit, a decision control unit and a simulation feedback unit. The decision optimization unit is configured to obtain a real-time optimal running trajectory of the target vehicle and a control parameter space under the difference scenario information by a vehicle trajectory control algorithm optimized by a particle swarm algorithm, according to the real-time running information of the target vehicle, the non-target vehicle running trajectory parameters under the difference scenario information, the pedestrian running trajectory parameters and the static obstacle distribution information within the visual range of the target vehicle. The decision control unit is configured to generate a real-time running control instruction of the target vehicle under the difference scenario information by combining a fuzzy control algorithm based on the real-time optimal running trajectory of the target vehicle and the control parameter space under the difference scenario information. The simulation feedback unit is configured to simulate the corresponding control abnormality and collision probability of the target vehicle under the generated real-time running control instruction by a simulation algorithm based on the running enhanced topological space of the target vehicle and the difference scenario information, and feed back the simulation result to the decision optimization unit to adjust the optimal running trajectory and control parameter space of the target vehicle in real time, so that the vehicle collision probability is minimized. 6.The deep learning based vehicle-mounted image recognition and target detection system of claim 5, wherein, The construction and training process of the multi-angle regional scene matching model includes: Extracting local visible boundary images and LiDAR point cloud distribution information of static environment distribution and static visible markers within the local visible boundary corresponding to the obstructed area from the target vehicle historical running records, and extracting multi-angle panoramic images and shooting angles of the obstructed area obtained by different non-target vehicles from the shared information corresponding to the same timestamp, as well as LiDAR point cloud distribution information of static environment distribution and static visible markers within the local visible boundary corresponding to the obstructed area in the panoramic image; According to the static environment distribution and the LiDAR point cloud distribution information of the static visible markers within the local visible boundary obtained by the target vehicle, and the static environment distribution and the LiDAR point cloud distribution information of the static visible markers within the local visible boundary corresponding to different non-target vehicles in the panoramic view at different viewing angles, the feature correspondence is performed through the timestamp, and the matching information pair is constructed through the aligned information; Based on the matching information pair, the first Mesh model of the local visible boundary corresponding to the occluded area is generated through the Poisson surface reconstruction algorithm, and the second Mesh model of the local visible boundary in the panoramic view corresponding to different viewing angles is generated, and the corresponding first modeling loss and second modeling loss are obtained; Based on the first Mesh model and the second Mesh model, automatic segmentation labeling is performed through the Panoptic-DeepLab model combined with the corresponding static marker labels, dynamic obstacle labels and traffic signal labels at the same time point within the local visible boundary, and the first Mesh model and the second Mesh model with complete labeling are obtained.

7. The deep learning-based in-vehicle image recognition and object detection system of claim 6, wherein, The construction and training process of the multi-angle regional scene matching model further includes: The first initial feature space and the second initial feature space are obtained by inputting the first Mesh model and the second Mesh model with complete labeling into the initial feature extraction layer in the multi-angle regional scene matching model; The first point cloud feature space and the second point cloud feature space and the first radar feature space and the second radar feature space are obtained by synchronously and parallelly inputting the first initial feature space and the second initial feature space into the point cloud feature extraction layer and the radar feature extraction layer, respectively; The multi-level multi-view feature fusion is performed through the multi-level feature fusion layer combined with the Bayesian neural network by inputting the first initial feature space and the second initial feature space, the first point cloud feature space and the second point cloud feature space, and the first radar feature space and the second radar feature space, to obtain the multi-level multi-view fusion feature space and the loss function corresponding to each level of feature fusion; The multi-level multi-view fusion feature space and the loss function corresponding to each level of feature fusion, the first modeling loss and the second modeling loss are input into the output layer combined with the cosine similarity function for training, and the matching accuracy of the local visible boundary of the target vehicle in the occluded area and the multi-view panoramic view corresponding to the non-target vehicle at each time is obtained, and the multi-angle regional scene matching model is trained. 8.The deep learning based vehicle-mounted image recognition and target detection system of claim 7, wherein, The construction process of the target vehicle operation enhancement topology space includes: The matching multi-view panoramic view corresponding to the same time point of the target vehicle in the occluded area is obtained according to the multi-angle regional scene matching model; Based on the first initial feature space and the second initial feature space, the first point cloud feature space and the second point cloud feature space, the first radar feature space and the second radar feature space corresponding to the target vehicle in the occluded area and the multi-view panoramic view, and the dynamic-static obstacle information and the traffic signal label information within the visible range of the target vehicle, the target vehicle enhanced three-dimensional space is obtained through the three-dimensional modeling algorithm combined with the map layer. Based on the target vehicle, the dynamic-static obstacle information and the traffic signal position information of each position point in the target vehicle's current time point corresponding visible range in three-dimensional space are enhanced, and the dynamic obstacle information and the traffic signal position information of each position point in the occluded area are constructed. Each time point target vehicle corresponding enhanced node space in the obstacle node and signal node; At the same time, based on the distance between the target vehicle and the static obstacle in the corresponding visible range at the same time point, a first connection relationship is constructed, and based on the distance between the target vehicle and the dynamic obstacle in the corresponding visible range at the same time point, the product of the collision probability and the probability of existing traffic signal label in the visible range is constructed. The second connection relationship is constructed. 9.The deep learning based vehicle-mounted image recognition and target detection system of claim 8, wherein, The construction process of the target vehicle running enhanced topological space also includes: And the product of the distance between the target vehicle and the vehicle in the corresponding occluded area at the same time point, the probability of collision, the probability of traffic signal existing in the occluded area, and the matching accuracy of the occluded area at the current time is constructed. The third connection relationship is constructed; At the same time, the product of the distance between the target vehicle and the pedestrian in the corresponding occluded area at the same time point, the prediction accuracy of the travel trajectory, the probability of collision, the probability of traffic signal existing in the occluded area, and the matching accuracy of the occluded area at the current time is constructed. The fourth connection relationship is constructed; Based on the obstacle node and signal node in the target vehicle corresponding enhanced node space at each time point, the corresponding first connection relationship, second connection relationship, third connection relationship and fourth connection relationship are combined, and the static target vehicle enhanced topological space corresponding to each time point is constructed by graph algorithm. 10.The deep learning based vehicle-mounted image recognition and target detection system of claim 9, wherein, The construction process of the target vehicle running enhanced topological space also includes: Based on the static target vehicle enhanced topological space, the vehicle trajectory analysis and prediction algorithm is configured on the second connection relationship and the third connection relationship, and the gait recognition algorithm is configured on the fourth connection relationship, combined with the multi-modal motion trajectory prediction algorithm and combined with the three-dimensional simulation algorithm. The dynamic running simulation of the target running vehicle in the traffic signal, no traffic signal and occluded and non-occluded area is carried out at continuous time points, and the target vehicle running enhanced topological space is obtained.

Citation Information

Patent Citations

  • An intelligent vehicle passable area detection method based on multi-source information fusion

    CN109829386A

  • Depth multi-mode perception unstructured scene automatic driving network architecture method

    CN118155183A