Abnormal parking detection method and device in expressway scene based on large model
By extracting features at multiple scales and processing environmental information, abnormal parking events in highway scenarios are identified, solving the accuracy problem of abnormal parking detection in highway scenarios and ensuring driving safety.
Patent Information
- Application Number
- CN202511285239.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-09
AI Technical Summary
In highway scenarios, existing technologies struggle to accurately identify abnormal parking events, impacting driving safety and efficiency.
A large model-based approach is adopted to identify the location and category of traffic elements and track their movement trajectories by extracting features at multiple scales, determining image quality information and environmental information, in order to identify abnormal parking events.
It improves the accuracy of traffic element recognition in complex scenarios, enabling timely detection of abnormal parking events and ensuring driving safety.
Smart Images

Figure CN120783300B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image processing, in particular to the technical field of artificial intelligence such as automatic driving, large language model, image feature detection, and more particularly to a large model-based abnormal parking detection method and device in a highway scene, an electronic device, a computer readable storage medium, and a computer program product. BACKGROUND
[0002] Compared with the urban road scene, the highway scene has the characteristics of fast driving speed and small external scene change, so the abnormal events in the highway scene are obviously different from those in the urban road scene, and the hazards caused by the abnormal events are also different.
[0003] Considering that the highway is an indispensable driving scene for long-distance travel, how to accurately identify abnormal events occurring in the highway scene and thus ensure traffic safety and efficiency is an important problem to be solved by those skilled in the art. SUMMARY
[0004] The embodiments of the present disclosure provide a large model-based abnormal parking detection method and device in a highway scene, an electronic device, a computer readable storage medium, and a computer program product.
[0005] In a first aspect, the embodiments of the present disclosure provide a large model-based abnormal parking detection method in a highway scene, comprising: acquiring a road time sequence video shot in a highway scene; performing a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain the multi-scale feature of each video frame; determining the image quality information and the environmental information of the corresponding video frame according to the multi-scale feature of each video frame, respectively; determining the position information and the category information of the traffic elements existing in the corresponding video frame according to the image quality information, the environmental information and the multi-scale feature of each video frame, respectively, wherein the image quality information and the environmental information are used to determine the confidence of the position and the category of the traffic elements; determining the motion trajectory of the same traffic elements according to the position information, the category information and the time sequence information between the video frames of the traffic elements existing in each video frame; and determining the abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic elements.
[0006] In a second aspect, the embodiments of the present disclosure provide an abnormal parking detection device based on a large model in a highway scene, comprising: a road time sequence video acquisition unit configured to acquire a road time sequence video shot in a highway scene; a multi-scale feature extraction unit configured to perform a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain multi-scale features of each video frame; an image quality and environment information determination unit configured to determine image quality information and environment information of each video frame according to the multi-scale features of each video frame, respectively; a traffic element recognition unit configured to determine position information and category information of a traffic element existing in each video frame according to the image quality information, the environment information and the multi-scale features of each video frame, respectively, wherein the image quality information and the environment information are used to determine the confidence of the position and the category of the traffic element; a motion trajectory determination unit configured to determine the motion trajectory of the same traffic element according to the position information, the category information and the time sequence information between each video frame of the traffic element existing in each video frame; and an abnormal parking event determination unit configured to determine an abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic element.
[0007] In a third aspect, the embodiments of the present disclosure provide an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the abnormal parking detection method based on a large model in a highway scene as described in the first aspect when executed.
[0008] In a fourth aspect, the embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the abnormal parking detection method based on a large model in a highway scene as described in the first aspect when executed.
[0009] In a fifth aspect, the embodiments of the present disclosure provide a computer program product comprising a computer program, which is used to enable a processor to implement each step of the abnormal parking detection method based on a large model in a highway scene as described in the first aspect when executed.
[0010] The abnormal parking detection scheme in the expressway scene based on the large model provided by the present disclosure first improves the recognition ability of traffic elements of different sizes in the video frame through the multi-scale feature extraction operation, then determines the image quality information and environmental information of the corresponding video frame according to the multi-scale features, and then determines whether the traffic elements reflected by the multi-scale features exist and the position and category of the traffic elements by means of the image quality information and environmental information of the video frame, so as to avoid the influence of image quality problems or environmental factors on the recognition accuracy of traffic elements, so that the motion trajectory based on the traffic elements identified can determine whether there is an abnormal parking event in the expressway scene even in a complex scene, and then the corresponding avoidance is performed in time to ensure the driving safety of the object in the expressway scene.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] Other features, objects, and advantages of the present disclosure will become more apparent through a detailed description of the non-limiting embodiments made by reading with reference to the following drawings:
[0013] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;
[0014] Figure 2 A flowchart of a large model-based abnormal parking detection method in an expressway scene provided by an embodiment of the present disclosure;
[0015] Figure 3 A flowchart of a method for obtaining image quality information and environmental information according to multi-scale features provided by an embodiment of the present disclosure;
[0016] Figure 4 A flowchart of a method for determining and obtaining optimized quality parameters and optimized environmental parameters according to learnable quality parameters and learnable environmental parameters provided by an embodiment of the present disclosure;
[0017] Figure 5 A flowchart of a method for determining an abnormal parking event according to a motion trajectory and for adjusting a large model provided by an embodiment of the present disclosure;
[0018] Figure 6a A flowchart of another large model-based abnormal parking detection method in an expressway scene provided by an embodiment of the present disclosure;
[0019] Figure 6b and Figure 6cA set of effect comparison charts of whether to apply the large model-based abnormal parking detection method in a highway scene provided in the present disclosure are provided for embodiments of the present disclosure;
[0020] Figure 6d and Figure 6e Another set of effect comparison charts of whether to apply the large model-based abnormal parking detection method in a highway scene provided in the present disclosure are provided for embodiments of the present disclosure;
[0021] Figure 7 A structural block diagram of a large model-based abnormal parking detection device in a highway scene provided for embodiments of the present disclosure;
[0022] Figure 8 A structural schematic diagram of an electronic device suitable for executing a large model-based abnormal parking detection method in a highway scene provided for embodiments of the present disclosure. DETAILED DESCRIPTION
[0023] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0024] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.
[0025] Figure 1 An exemplary system architecture 100 that can apply embodiments of the large model-based abnormal parking detection method, device, electronic device and computer readable storage medium in a highway scene of the present disclosure is shown.
[0026] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a communication link medium between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0027] The user can use the terminal device 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications for realizing information communication between the terminal device 101, 102, 103 and the server 105 can be installed on the terminal device 101, 102, 103 and the server 105, such as image processing applications, abnormal event detection applications, model training applications, etc.
[0028] The terminal device 101, 102, 103 and the server 105 can be hardware or software. When the terminal device 101, 102, 103 is hardware, it can be various electronic devices with a display screen, including but not limited to roadside cameras, high-speed monitoring cameras, car camera cameras, etc.; when the terminal device 101, 102, 103 is software, it can be installed in the above-mentioned electronic devices, which can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; when the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here.
[0029] The server 105 can provide various services through various built-in applications. Taking an abnormal event detection application that can provide a video-based abnormal parking event detection service as an example, the server 105 can achieve the following effects when running the electronic album application: first, acquire a road time sequence video obtained by shooting a highway scene through the terminal device 101, 102, 103; then, perform a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain the multi-scale feature of each video frame; next, determine the image quality information and environmental information of each video frame according to the multi-scale feature of each video frame, respectively; and determine the position information and category information of the traffic elements existing in the corresponding video frame according to the image quality information, environmental information and multi-scale feature of each video frame, respectively. The image quality information and the environmental information are used to determine the confidence of the position and category to which the traffic elements belong; next, determine the motion trajectory of the same traffic elements according to the position information, category information and time sequence information between the video frames; finally, determine the abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic elements.
[0030] It should be noted that in addition to being acquired from the terminal devices 101, 102, 103 through the network 104, the road time sequence videos can also be pre-stored locally in the server 105 in various ways. Therefore, when the server 105 detects that the data has been stored locally (for example, before starting to handle the remaining to-be-handled detection tasks), it can be selected to directly acquire the data locally, and in this case, the exemplary system architecture 100 can also not include the terminal devices 101, 102, 103 and the network 104.
[0031] Since detecting abnormal parking events based on road time sequence videos requires occupying more computing resources and stronger computing capabilities, the abnormal parking detection method in the high-speed road scene based on a large model provided by each of the subsequent embodiments of the present disclosure is generally executed by the server 105 with stronger computing capabilities and more computing resources. Correspondingly, the abnormal parking detection device in the high-speed road scene based on a large model is also generally arranged in the server 105. However, it should also be noted that when the terminal devices 101, 102, 103 also have computing capabilities and computing resources that meet the requirements, the terminal devices 101, 102, 103 can also complete the above-mentioned operations by the server 105 through the abnormal event detection application installed thereon, and then output the same results as the server 105. Especially in the case where there are multiple terminal devices with different computing capabilities, but the abnormal event detection application judges that the terminal device has strong computing capabilities and has more remaining computing resources, the terminal device can be allowed to execute the above-mentioned operations, thereby appropriately reducing the computing pressure of the server 105. Correspondingly, the abnormal parking detection device in the high-speed road scene based on a large model can also be arranged in the terminal devices 101, 102, 103. In this case, the exemplary system architecture 100 can also not include the server 105 and the network 104.
[0032] It should be understood that Figure 1 The number of terminal devices, networks and servers in
[0033] Reference is made to Figure 2 , Figure 2 A flowchart of an abnormal parking detection method in a high-speed road scene based on a large model provided by an embodiment of the present disclosure, wherein the flow 200 includes the following steps:
[0034] Step 201: acquiring road time sequence videos shot in a high-speed road scene;
[0035] This step aims to acquire road time sequence videos by the execution subject of the abnormal parking detection method in the high-speed road scene based on a large model (for example Figure 1The server 105 shown) obtains a road time-series video taken by a highway scene.
[0036] In the highway scene, the collection of video is usually realized by fixedly installed high-definition monitoring cameras or mobile patrol vehicles equipped with car recorders. These devices should generally have appropriate focal lengths and viewing angles to ensure coverage of key monitoring areas (such as lanes, emergency lanes, and ramp entrances). Considering the dynamic characteristics of the highway scene, the video collection device should also be able to capture the details of fast-moving vehicles at a high frame rate (usually not less than 25 fps), while maintaining sufficient image resolution (1080p or higher) to support multi-scale feature extraction used in subsequent steps.
[0037] In actual operation, the video data can be transmitted in real time to an edge computing device or a cloud processing center through a wired network or a 5G wireless network. In this process, data compression processing (such as encoding according to the H.264 / H.265 encoding method) can also be performed to balance bandwidth and image quality, and time stamps and location information (associated through GPS or camera ID) can also be embedded to provide a spatio-temporal context for subsequent time-series analysis. For the obtained original video stream, especially the multi-channel video stream obtained from different devices and different acquisition channels, the above execution subject can also perform preliminary standardized preprocessing, including but not limited to format unification (such as converting to a standard RGB sequence), basic denoising (for particle noise in rainy and snowy weather), and automatic exposure / white balance adjustment (to cope with light mutations at tunnel entrances and exits), etc.
[0038] In the case that the target highway scene corresponds to multiple road videos with different perspectives, a specific and non-limiting processing method can be: performing perspective fusion on the multiple road videos to obtain a road time sequence video corresponding to the target highway scene after perspective completion. Specifically, when the target highway section has a heterogeneous visual network composed of multiple monitoring cameras (such as main line card hole cameras, roadside wide-angle cameras, overhead cameras, etc.), the above execution subject can first perform spatio-temporal alignment on each video source, for example, by using pre-calibrated camera internal and external parameters (such as focal length, pitch angle, installation height) and on-site measured geographic location information to establish the spatial projection relationship between perspectives; at the same time, millisecond-level time synchronization is realized by using network time protocol or video frame timestamp, so as to eliminate the time sequence deviation caused by independent triggering of the camera. In the specific fusion process, the above execution subject can also use a multi-perspective geometric completion algorithm to jointly model the local scene captured by different cameras (such as vehicle front features under the main perspective, lane line extension under the side perspective, and vehicle flow distribution under the overhead perspective): based on feature matching (such as lane line vanishing point alignment and road texture correlation), a coordinate mapping relationship across perspectives is established, and a virtual bird's eye perspective is generated through three-dimensional space interpolation, so as to eliminate the visual blind area (such as large vehicle shielding area and ramp merging place) under the traditional single perspective. For the fusion of dynamic targets, the above execution subject can also combine the confidence of the detection results of each perspective (such as giving a higher weight to the main perspective due to the short distance) to associate the targets across perspectives, and try to use Kalman filtering to predict the projection consistency of the trajectory under multiple perspectives, so as to solve the target ID jump problem caused by perspective switching.
[0039] In this case, the finally generated completed time sequence video not only contains the spatio-temporal synchronization data stream of the original multi-channel video, but also integrates the complete road topology (such as the number of lanes and the position of the separation strip) and the three-dimensional distribution of dynamic elements reconstructed by perspective fusion. The panoramic scene expression will enable the subsequent abnormal parking detection to avoid the limitations of single perspective (such as long-distance small target missed detection and shielding caused trajectory interruption), and significantly improve the robustness of event detection under complex road sections. Further, in actual engineering implementation, the influence of the optical property differences (such as color difference and distortion degree) of different cameras on the fusion effect can be considered, and light consistency correction and dynamic white balance adjustment can be introduced to ensure the visual continuity of multi-source videos.
[0040] Step 202: performing a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain multi-scale features of each video frame;
[0041] On the basis of step 201, this step aims to perform a multi-scale feature extraction operation on each video frame constituting the road time sequence video by the above-mentioned execution subject, to obtain the multi-scale features of each video frame, so as to better cope with the richness and diversity of traffic elements in the scale level in the high-speed scene through hierarchical visual feature capture.
[0042] At the specific practice level, the above-mentioned execution subject can adopt a feature pyramid architecture based on a deep convolutional neural network to perform multi-level semantic abstraction on the input video frame: in the shallow convolution stage, dense convolution kernels with small receptive fields are used to extract pixel-level local features (such as edge texture, color gradient), which are crucial for identifying the details of large vehicles in the near distance (such as license plate, car light arrangement); with the increase of network depth, regional features (such as wheel contour, window shape) are captured in the middle layer feature map through gradually expanding receptive field and down-sampling operation, which are suitable for component-level identification of vehicles at medium distance; and in the deep network stage, low-resolution but rich global semantic high-order features (such as vehicle overall appearance, orientation angle) are generated, which are specially used for processing the category determination of small-size vehicles at a distance.
[0043] In order to keep the feature response of small targets from being diluted by deep networks, the above-mentioned execution subject can also use a cross-level feature fusion mechanism to perform element-by-element superposition of high-resolution features in shallow layers and high semantic features in deep layers through up-sampling, forming a hybrid expression with details and semantics. For the long-distance imaging characteristics of expressways, this process can also incorporate an adaptive scale enhancement strategy to dynamically adjust the feature extraction granularity of different regions according to the perspective relationship of the current frame (such as the position of the vanishing point): more dense small-scale feature extraction is used for the distant congested traffic area, while medium and large-scale feature mining is focused on the near sparse vehicle area. The finally generated multi-scale feature tensor actually constitutes a microscopic to macroscopic visual description system, each spatial position is associated with a feature vector of different abstraction degree, these vectors not only encode the morphological information (such as size, shape) of traffic elements, but also imply the context relationship (such as the relative position of vehicles and lane lines) with the surrounding environment, providing a granular rich basic feature library for subsequent robust detection of image quality and environmental information fusion.
[0044] Further, considering that the computational efficiency of this module directly affects the real-time performance of the system in actual deployment, a lightweight backbone network can also be used in combination with pruning and quantization techniques to ensure multi-scale expression capability while meeting the timeliness requirements of video stream processing. It should be noted that the model mentioned above that can realize multi-scale feature extraction can exist as part of the overall abnormal event detection large model or as an external support model of the large model.
[0045] Step 203: determining image quality information and environment information of each video frame according to multi-scale features of each video frame, respectively;
[0046] On the basis of step 202, this step aims to determine image quality information and environment information of each video frame according to multi-scale features of each video frame by the above execution subject, so as to realize accurate evaluation of scene objective conditions by deeply mining multi-level visual features of video frames. The environment information includes weather sub-information and illumination intensity sub-information, the weather sub-information includes at least one of normal (mainly refers to sunny day), rain, wind, fog, haze, thunder and lightning, and snow, and the illumination intensity sub-information includes at least one of normal, too bright, and too dark. The image quality information can include at least one of normal (i.e. without various quality problems mentioned below), screen flower, trailing, ghosting, overexposure, overdarkness, lower than preset resolution, and noise quantity exceeding preset quantity.
[0047] At a specific practice level, the above execution subject can first utilize response modes of different levels in multi-scale features to identify image quality problems: shallow high-frequency texture features (such as edge sharpness) can detect screen flower, trailing and ghosting phenomenon, when these features abnormally fluctuate or directionally disorder in local area, it indicates that there may be image transmission error or motion blur; mid-layer contrast features combined with color distribution can judge overexposure or overdarkness, such as large area high brightness region accompanied by color saturation drop indicating overexposure, while low contrast and concentrated in dark tone pointing to overdarkness; deep layer semantic features are used to identify global quality problems, such as when detecting a clear license plate region but presenting a semantic contradiction of blur, it can be determined that the resolution is insufficient. For noise detection, the above execution subject can analyze the response strength of multi-scale features in flat area, and abnormal high-frequency noise is manifested as irregular fine-grained feature activation.
[0048] And in the aspect of environment information analysis, the recognition of weather sub-information can rely on joint analysis of multi-scale features: rain and snow weather is manifested as dense vertical lines (trajectories of raindrops / snowflakes) in shallow features and transmittance drop (contrast reduction caused by fog) in mid-layer features; lightning weather is captured by capturing instantaneous brightness mutation features; and strong wind weather needs to combine with motion trajectory anomaly of dynamic targets (such as tree swaying) to judge. The determination of illumination intensity sub-information can mainly depend on global statistics of mid-layer features: under normal illumination, feature response is evenly distributed; when too bright, high brightness area features dominate and are accompanied by missing shadow area features; when too dark, overall feature response strength is low.
[0049] The above-mentioned quality and environment evaluation results can be output in the form of a multi-dimensional vector, each dimension corresponding to a confidence score under a specific condition. Subsequent stages will dynamically adjust the detection strategy accordingly, such as increasing the weight of motion features in rainy and foggy weather, or suppressing the output of color information-dependent classifiers in overexposed conditions, thereby ensuring stable abnormal parking detection performance in various complex scenarios.
[0050] Step 204: Determine the position information and category information of the traffic elements present in each video frame according to the image quality information, environmental information, and multi-scale features of each video frame, respectively;
[0051] On the basis of step 203, this step aims to determine the position information and category information of the traffic elements present in each video frame according to the image quality information, environmental information, and multi-scale features of each video frame, respectively, by the above-mentioned execution subject, in order to accurately determine the position information and category information of the traffic elements present in each video frame. That is, this step aims to dynamically integrate image quality evaluation, environmental perception, and multi-scale feature expression to construct a scene-adaptive detection mechanism.
[0052] At a practical level, the above execution subject can adopt a multi-branch collaborative reasoning architecture: a multi-scale feature-based backbone network first generates an initial traffic element candidate box and its coarse-grained location (bounding box coordinates) and class (vehicle, pedestrian, obstacle, etc.) prediction, and then image quality information and environmental information are introduced as modulation factors in the refinement process. For location information, the above execution subject can implement spatial confidence weighting according to the image quality evaluation results: in areas where ghosting or ghosting occurs, the weight of location regression is reduced to avoid positioning drift caused by blurred edges; for low-resolution or overexposed areas, the bounding box coordinates are corrected by motion consistency compensation of adjacent frames. In terms of class determination, a feature selection strategy based on environmental information guidance can be used: in rainy and foggy weather, the contribution of contour shape features is strengthened (because color features are unreliable), and in strong light conditions, emphasis is placed on the analysis of shadow and texture features. At the same time, the above execution subject can also establish a condition-dependent confidence decay model: when extreme weather (such as heavy rain) or severe image degradation (such as severe noise) is detected, the confidence threshold of the detection results in the corresponding area is automatically lowered to avoid misjudgment caused by low-quality data. This fusion mechanism finally outputs the precise location (spatial coordinates with uncertainty range), fine-grained class (such as car / truck / hazardous vehicle), and environment-calibrated confidence score of each traffic element, which not only reflects the model's grasp of the target's existence, but also contains a quantitative assessment of the positioning accuracy and classification reliability. For example, a truck identified in a thin fog environment may have a high existence confidence (due to its obvious volume feature), but the location confidence will be moderately lowered due to visibility. These reliability-labeled detection results provide quality-controllable input for subsequent trajectory tracking, ensuring that abnormal parking judgments are neither overly conservative and miss real events, nor produce false positives due to noise interference.
[0053] In addition to the above-mentioned specific implementation, another implementation can be: first, determine the confidence of each video frame according to the image quality information and environmental information of each video frame, respectively; then only identify the location information and class information of the existing traffic elements from the video frames with a confidence exceeding the pre-set confidence threshold. This approach aims to effectively filter out the interference of low-quality data on subsequent analysis by quantitatively evaluating the availability of single-frame images.
[0054] At the level of specific practice, the above execution subject can first build a multi-dimensional confidence evaluation model: for image quality information, the degradation types such as screen flicker and trailing are mapped to 0-1 quantization scores according to the severity (such as local screen flicker deducts 0.3, and serious global trailing deducts 0.7), and the spatial overlap rate of the degradation area and the target detection area is combined for weighted calculation; for environmental information, a weather-light coupling score matrix is established (such as the confidence base value of "heavy rain + night" combination is 0.2, and the "sunny + normal light" is 0.9), while considering the time persistence of environmental factors (continuous multiple frames of bad weather need to be additionally reduced threshold). The above execution subject can also dynamically calculate the comprehensive confidence index, which not only includes the above objective evaluation, but also introduces the quality stability analysis of historical frames: if the confidence suddenly drops in the current frame, but the previous and next frames are normal, it may trigger temporary data storage instead of direct discard. For qualified video frames that exceed the preset threshold (usually set in the interval of 0.5-0.7), the above execution subject can also start the traffic element recognition of the whole process, at this time the confidence score will be embedded in the detection process as metadata: in the target positioning stage, high-confidence frames allow the use of a more lenient non-maximum suppression threshold to retain more candidate boxes, while low-confidence frames need to strictly verify the spatial consistency of the bounding box; in the classification stage, the confidence level determines the feature fusion strategy, and high-confidence frames can enable more computationally expensive fine-grained classifiers (such as distinguishing between ambulances and ordinary vans). Video frames judged as low confidence will enter the cross-frame compensation processing flow, generate virtual detection results through motion prediction of adjacent high-confidence frames, but will be marked with their speculative properties to avoid misleading trajectory analysis.
[0055] Such strict quality gating mechanism will enable the maintenance of basic vehicle detection capability based on the most reliable data when encountering extreme weather or equipment failure, while suspending the judgment of abnormal parking and other fine events, thereby maintaining a reasonable lower limit of performance under complex environmental fluctuations. At the same time, all confidence evaluation parameters support online learning adjustment, by continuously monitoring the correlation between false positives / misses and confidence, dynamically optimizing the threshold setting to adapt to the particularity of specific road sections (such as foggy road sections can appropriately adjust the fog deduction weight).
[0056] Step 205: determining the motion trajectory of the same traffic element according to the corresponding position information, category information and time sequence information between the video frames;
[0057] On the basis of step 204, this step aims to determine the motion trajectory of the same traffic element according to the corresponding position information, category information and time sequence information between the video frames by the above execution subject. This process matches the same traffic elements across frames through a spatio-temporal correlation algorithm, thereby realizing the tracking of motion trajectory.
[0058] To achieve the purpose of this step, cross-frame target association can be performed first, such as optimal matching of detection boxes of adjacent frames using the Hungarian algorithm or graph neural networks, and the matching basis includes not only traditional appearance similarity and motion consistency (such as the overlap rate of the predicted position by Kalman filtering and the actual detection IoU), but also the class confidence and environment adaptation factor from the previous module: increasing the weight of motion features in rainy and foggy weather, and strengthening the feature comparison of vehicle lights in low light conditions; secondly, a trajectory segment optimization mechanism can be constructed, which uses a smoothing algorithm with constraints (such as quadratic programming considering lane line constraints) to correct trajectory drift caused by temporary occlusion or detection jitter, and introduces speed-acceleration physical rules to verify the rationality of the trajectory and filter abnormal connections that obviously violate the dynamics of the vehicle; finally, full-cycle trajectory fusion can be implemented, which solves the long-term occlusion problem (such as a large truck occluding a small car) through multiple hypothesis testing in a sliding time window, and when the target reappears, it comprehensively considers the position extrapolation result before disappearance, the possible driving path (based on lane topology constraints), and the feature matching degree when it reappears to confirm the identity. The above execution subject maintains a dynamically updated trajectory descriptor for each confirmed traffic target, including spatiotemporal position sequence, motion state (speed / acceleration), appearance feature codebook, and environment-dependent reliability markers. These rich contextual information can enable the subsequent abnormal parking judgment to distinguish between active braking (such as avoiding an oncoming vehicle) and abnormal stasis (such as a malfunctioning vehicle), while effectively resisting trajectory breaks caused by temporary occlusion or detection failure.
[0059] Further, considering that in actual engineering deployment, the problem of computational complexity under extremely dense traffic flow often needs to be addressed, a ROI (Region of Interest) based partition tracking strategy is usually used to decompose the global association problem into lane-level sub-problems, and the ID switching problem of vehicles changing lanes is solved through trajectory conflict detection and arbitration mechanism, ensuring that the trajectory integrity is maintained at >95% even under the traffic flow of hundreds of vehicles per minute.
[0060] Step 206: determining the abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic element.
[0061] On the basis of step 205, this step aims to determine the abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic element by the above execution subject. That is, to convert continuous trajectory data into decision-making judgments that conform to traffic management semantics.
[0062] To achieve this purpose, it is necessary to realize precise event detection through multi-dimensional spatio-temporal feature analysis, which can be specifically broken down into the following levels of judgment logic: First, motion state analysis is performed, the speed-time curve and acceleration-displacement distribution of each vehicle trajectory are calculated, and the speed drop point is identified through a dynamic threshold algorithm (such as a 2-second drop from 100 km / h to 10 km / h), while combining high-precision map data to exclude allowed deceleration areas (such as legal deceleration within 1 km before a toll station); second, parking position verification is implemented, using pre-labeled lane semantic information (driving lane, overtaking lane, emergency lane) and real-time traffic flow state (average speed of surrounding vehicles) to determine whether the stopping position violates traffic rules: stopping for more than 10 seconds in a non-emergency lane triggers a primary alarm, while stopping in high-risk areas such as curves or tunnels uses a stricter 5-second threshold; third, behavior context correlation analysis is performed, using trajectory cluster detection technology to identify multi-vehicle correlation behaviors (such as chain parking caused by a front accident), and to distinguish between active parking (courtesy for an ambulance) and passive abnormalities (mechanical failure), at which point dynamic baselines are established using historical data, such as temporary slowdowns on the same road during the morning rush hour being normal, while sudden stops during off-peak hours may indicate an event; finally, multi-evidence fusion decision-making is performed, inputting motion features (stopping time), spatial features (lane deviation), environmental factors (visibility impact in rainy or foggy weather), and device confidence (image quality score of the detected vehicle) into a fuzzy reasoning system, outputting an abnormal parking event report with probability assessment, including event type (failure / accident / violation parking), danger level (classified as 1-3 levels based on stopping position and duration), and recommended handling measures (such as immediate police dispatch or only monitoring and observation).
[0063] Based on the above, the executing subject can also continuously monitor the subsequent state changes of the identified abnormal parking target (such as turning on the double flasher, personnel getting off, etc.), and dynamically adjust the response strategy by establishing an event lifecycle model, ensuring a complete closed loop from detection to disposal. In practical applications, it is also possible to try to link with external systems such as weather data and real-time traffic, for example, automatically increasing the monitoring sensitivity of emergency lane parking during heavy rain warnings, or temporarily relaxing the time threshold for parking judgment on this road when a front accident is known, reflecting the situational awareness and adaptive ability of intelligent transportation systems.
[0064] The method for detecting abnormal parking in a highway scene based on a large model provided by the embodiments of the present disclosure first improves the recognition ability of traffic elements of different sizes in the video frame through multi-scale feature extraction operation, and then determines the image quality information and environmental information of the corresponding video frame according to the multi-scale features, and then determines whether the traffic elements reflected by the multi-scale features exist and the position and category of the traffic elements by means of the image quality information and environmental information of the video frame, so as to avoid the influence of image quality problems or environmental factors on the recognition accuracy of traffic elements, so that the motion trajectory based on the traffic elements identified can determine whether there is an abnormal parking event in the highway scene even in a complex scene, and then the corresponding avoidance is performed in time to ensure the driving safety of the object in the highway scene.
[0065] To further understand how step 203 determines the image quality information and environmental information based on the multi-scale features, please refer to Figure 3 , Figure 3 A flowchart of a method for obtaining image quality information and environmental information according to multi-scale features provided by the embodiments of the present disclosure, wherein the flowchart 300 includes the following steps:
[0066] Step 301: Perform feature fusion processing on the multi-scale features of each video frame respectively to obtain multi-scale fusion features of each video frame;
[0067] To realize the fusion of multi-scale features, a channel-space dual combination strategy can be specifically used, for example, the contribution of different scale features can be weighted by a channel attention mechanism first, and the feature maps of each layer of the pyramid are dynamically fused according to semantic importance; at the same time, cross-scale feature alignment is implemented in the spatial dimension, and deformable convolution is used to compensate for the geometric deformation caused by scale changes, and finally a unified feature representation containing global context and local details is generated. It should be noted that this fusion method is adapted to the characteristics of near and far targets coexisting in the highway scene, which can ensure that the fine structure of the near vehicle and the contour information of the far vehicle can be retained.
[0068] A specific implementation mode can be: for each video frame, the following feature fusion processing operation is performed: average pooling processing is performed on the original features of each scale of the video frame to obtain length-unified features of each scale; the length-unified features of each scale of the video frame are spliced, and the features obtained after splicing are determined as the multi-scale fusion features of the corresponding video frame.
[0069] The present implementation attempts to construct a unified feature space through standardization processing and dimension reduction. First, scale normalization processing is performed. Global average pooling operations are applied to each scale branch (such as the feature maps output by different stages of the residual network) in the original multi-scale feature map, respectively. Different resolution feature maps (such as small-scale high-level semantic features of 56x56 and large-scale detailed features of 224x224) are compressed into fixed-length vector representations. This process essentially retains the global response intensity of each feature channel while ignoring spatial distribution differences. It is particularly suitable for scene understanding tasks in highway scenarios that need to emphasize the overall properties of the target rather than the precise location. Then, cross-scale feature alignment can be performed. According to the hierarchical depth of each scale feature in the backbone network, a learnable linear transformation is performed on the pooled feature vector to compensate for the feature distribution offset caused by different network depths (such as the difference between the sparsity of deep features and the density of shallow features). This ensures that features of different abstraction levels can be in a comparable numerical interval. Finally, a splicing fusion is implemented. The normalized and aligned scale feature vectors are concatenated in the channel dimension to form a composite feature representation with multi-level semantics. Although this splicing strategy may sacrifice some spatial correspondence, it can maximize the avoidance of the small target feature submersion problem caused by traditional weighted fusion (such as the features of a distant vehicle not being covered by the strong response of a nearby large vehicle) by preserving the independence between original scales.
[0070] In practical applications, this scheme is particularly suitable for handling the cross-lane multi-target scene commonly seen in highway monitoring. The global nature of average pooling can balance the capture of vehicle information at different distances, and simple splicing operations ensure that the system can still maintain real-time processing capability under limited computing resources. Although the fused features lose fine-grained spatial information, they can still effectively support image quality and environmental information discrimination analysis through subsequent learnable parameter attention mechanism compensation (such as the cross-attention operation described in step 303). This design will demonstrate its advantages when deployed on edge computing devices.
[0071] Step 302: Initialize learnable quality parameters and learnable environmental parameters corresponding to the respective video frames with the same size and dimension number as the multi-scale fusion feature dimension number.
[0072] On the basis of step 301, the essence of the learnable parameter initialization process described in this step is to construct an auxiliary learning carrier matched with the feature space, and the dimension design is strictly corresponding to the channel number of the multi-scale fusion feature. Among them, the quality parameter and the environment parameter are respectively initialized as two independent high-dimensional vectors, and each dimension of each vector is associated with a specific semantic concept of the feature space (such as "flower screen", "ghosting" quality dimension or "weather", "light intensity" environment dimension). These parameters are initialized with small random numbers at the beginning of training, and gradually form the decomposition expression ability of complex scenes through subsequent learning.
[0073] Step 303: Cross attention processing of learnable quality parameters and learnable environment parameters with multi-scale fusion features of the same video frame to obtain learned quality parameters and learned environment parameters;
[0074] On the basis of step 302, the cross attention processing described in this step aims to realize the interactive enhancement of multi-scale fusion features and learnable parameters, and the core is to establish a bidirectional information flow mechanism: on the one hand, the multi-scale fusion features are used as Key and Value, and the feature area most relevant to quality / environment judgment is selected through the query-key value matching mechanism; on the other hand, the learnable parameters are used as Query, and the abstract concept is injected into the feature space to guide the network to focus on the feature dimension with discriminability. This cross action will produce parameter update of scene perception, such as automatically strengthening the "rain" dimension weight in the environment parameter when detecting raindrop-like texture features.
[0075] Step 304: Self-attention processing of learned quality parameters and learned environment parameters to obtain optimized quality parameters and optimized environment parameters;
[0076] On the basis of step 303, the self-attention processing described in this step is a deep mining of the internal relationship of learnable parameters, which forms the collaborative optimization between parameters by calculating the correlation between quality parameter dimensions (such as the co-occurrence mode of "ghosting" and "ghosting") and the coupling relationship between environment parameter sub-information (such as the accompanying probability of "heavy rain" and "low light"). The processing can adopt a multi-head attention architecture to capture long-range dependencies between parameters from different subspaces, which is particularly suitable for processing complex scene degradation (such as multi-factor interference of fog and lens pollution) commonly seen in highway monitoring.
[0077] Step 305: Nonlinear transformation processing and classification prediction processing are performed on the optimized quality parameters and optimized environment parameters of each video frame in turn to obtain the image quality information and environment information of the corresponding video frame.
[0078] On the basis of step 304, the step first performs nonlinear mapping on the optimized parameters, converts a high-dimensional vector into a more discriminative hidden space representation, and then can adopt a multi-branch classification head structure to output quality and environment information in parallel, wherein the quality classifier adopts a multi-label softmax processing (allowing the coexistence of "overexposure" and "noise" complex conditions), and the environment classifier is designed as a hierarchical structure (judging the "weather" category first and then subdividing the "rain / snow" subcategory).
[0079] A specific implementation can be: for each video frame, the following multi-layer perception processing is performed: the optimized quality parameter and the optimized environment parameter of the video frame are sequentially processed through the first full connection layer and the nonlinear transformation layer respectively to obtain the processed quality feature and the processed environment feature; wherein the first full connection layer is used to increase the feature dimension of the optimized quality parameter and the optimized environment parameter; the processed quality feature and the processed environment feature of the video frame are sequentially processed through the second full connection layer and the classification prediction layer respectively to obtain the image quality information containing the image quality category and the probability size and the environment information containing the environment category and the probability size; wherein the second full connection layer is used to reduce the feature dimension of the processed quality feature and the processed environment feature.
[0080] The implementation completes the decoding process of high-level semantic information through two-stage dimension regulation: the first stage implements feature space expansion, inputs the optimized quality parameter and the environment parameter into the first full connection layer with expansion effect, constructs a high-dimensional hidden space by increasing the feature dimension (such as from the original 256 dimensions to 1024 dimensions), and this dimension expansion operation essentially provides more combination possibilities for subsequent nonlinear transformation, which is particularly beneficial to decoupling complex scene features (such as processing "rain and fog + lens contamination" complex quality degradation). Nonlinear transformation (such as GELU activation function) is performed in the expanded high-dimensional space, and the multi-order derivative characteristics can better model the nonlinear relationship in the quality and environment parameters (such as the exponential relationship between the light intensity and the image noise). The second stage projects the high-dimensional features into the low-dimensional classification space required by the task by using the second full connection layer. This structure design of expanding the dimension first and then compressing the dimension is equivalent to constructing a "learning bottleneck" in the feature flow, which can force the network to learn more discriminative feature representations in the previous stage. The final classification prediction layer adopts a multi-task adaptive architecture: for the image quality information, a softmax processing with a temperature coefficient is used to control the prediction sharpness of the co-occurrence problems such as blur / overexposure by adjusting the temperature parameter; for the environment information, a hierarchical classifier structure is adopted, first performing binary coarse classification of "severe / normal", and then subdividing the specific types (such as heavy rain / snow) of the confirmed severe environment, which ensures the classification accuracy of complex scenes and avoids the decision ambiguity caused by too many categories.
[0081] The entire processing flow realizes dynamic balance of feature expression and classification requirements through dimension scaling, while maintaining lightweight models, accurately identifies various subtle quality degradation patterns (such as progressive image noise under low illumination at night) and complex environmental conditions (such as fog accompanied by lateral strong light) in highway monitoring, and provides reliable scene perception prior for subsequent traffic element detection. The scheme shows good precision-efficiency balance when deployed on edge devices, and can flexibly adapt to the application requirements of different computing platforms by adjusting the dimension scaling ratio of the fully connected layer.
[0082] The embodiment provides, through steps 301-305, a joint optimization of feature extraction and information prediction through end-to-end training. The final output quality and environment labels not only serve as confidence basis for subsequent detection modules, but also are fed back to the feature fusion stage to form a closed-loop optimization, and have self-adaptive analysis capability for complex scenes.
[0083] For further understanding of the cross-attention processing and self-attention processing provided by steps 303 and 304 respectively, please refer to Figure 4 , Figure 4 A flowchart of a method for determining optimized quality parameters and optimized environment parameters according to learnable quality parameters and learnable environment parameters provided by the embodiment of the disclosure, the flow 400 includes the following steps:
[0084] Step 401: Perform cross-attention calculation by taking the learnable quality parameters of the same video frame as the query information, and taking the multi-scale fusion features as the key information and the value information, to obtain learned quality parameters that learn image quality related information from the multi-scale fusion features;
[0085] This step provides a specific implementation of using a query-key-value separated attention mechanism to extract quality related features, that is, taking the learnable quality parameters as a dynamic query vector to perform semantic retrieval in the feature library composed of multi-scale fusion features. In specific implementation, the query vector calculates the similarity score (through dot product attention or cosine similarity) with each spatial position of the feature map, and focuses on those feature regions highly related to typical quality degradation patterns (such as enhancing the attention to time high frequency components when motion blur is detected). This mechanism enables the above execution subject to adaptively adjust the quality evaluation focus according to the current frame characteristics, for example, automatically enhancing the sensitivity to transient overexposure features in the tunnel entrance scene.
[0086] Step 402: Perform cross-attention calculation by taking the learnable environment parameters of the same video frame as the query information, and taking the multi-scale fusion features as the key information and the value information, to obtain learned environment parameters that learn environment related information from the multi-scale fusion features;
[0087] This step is similar to step 401, the difference is that this step is used to extract environment-related features.
[0088] Step 403: Determine the first target feature in the learned quality parameter that has the highest correlation with image quality, and determine the correlation between the features in the learned quality parameter in a self-attention processing manner, and eliminate other features that have an actual correlation less than a first preset correlation threshold with the first target feature, to obtain an optimized quality parameter.
[0089] This step aims to optimize the internal features through the self-attention mechanism. The above execution subject can first identify the feature dimension in the learned quality parameter that has the strongest statistical correlation with typical image quality problems (such as blur, overexposure, etc.) through a correlation score matrix (such as a similarity measure based on attention score), and mark it as the first target feature. For example, when analyzing tunnel scene video frames, the above execution subject may find that a certain group of feature dimensions are the most indicative of "instantaneous overexposure", and the activation pattern of these dimensions is highly synchronized with the light mutation when the vehicle enters / leaves the tunnel. Subsequently, a multi-head self-attention mechanism can be further used to construct a feature dependency graph to quantify the cooperative relationship between each dimension and the first target feature, for example, identifying that the "dynamic blur" feature has a negative correlation with the "overexposure" feature (because fast exposure adjustment causes motion blur), while the "noise" feature has a weak correlation with "overexposure". The system implements feature pruning according to a preset dynamic threshold (initial value set to 0.3-0.5 interval, which can be adjusted adaptively according to the scene), and retains those feature combinations that form stable prediction patterns with core quality problems, and eliminates redundant dimensions with low contribution. This optimization allows the final quality parameter to retain the ability to distinguish key quality degradation patterns while avoiding information overlap between dimensions.
[0090] Step 404: Determine the second target feature in the learned environment parameter that has the highest correlation with the environment, and determine the correlation between the features in the learned environment parameter in a self-attention processing manner, and eliminate other features that have an actual correlation less than a second preset correlation threshold with the second target feature, to obtain an optimized environment parameter.
[0091] Different from step 403, this step is used for optimizing the environmental characteristics. When identifying the second target feature most related to the environmental factors (such as rain, fog, and light), a spatiotemporal consistency verification can also be introduced, for example, requiring the "rain intensity" related feature to show progressive enhancement in continuous multiple frames, and the spatial distribution to conform to the physical characteristics of rain (wet diffusion from top to bottom). In the self-attention stage, the environmental parameter optimization pays special attention to the causal relationship between features, for example, distinguishing whether "low light" is caused by a night environment or tunnel shielding, which is achieved by analyzing the duration of feature activation and the coordinated change of the surrounding vehicle light state. During the optimization process, the above execution subject can also maintain an environmental feature importance ranking, dynamically increasing the feature dimension weight of those that have recently contributed to the judgment of complex environments (such as increasing the decision weight of the "rain line density" feature during sudden heavy rain), forming an adaptive parameter system with short-term memory capability.
[0092] The embodiment gives a specific implementation of step 303 through steps 401-402, and a specific implementation of step 304 through steps 403-404. The two specific implementations do not have a causal and dependent relationship, and can completely form different embodiments in the form of an alternative upper scheme. This example only exists as a preferred embodiment at the same time.
[0093] On the basis of any of the above embodiments, to deepen the understanding of how to determine an abnormal parking event based on a motion trajectory and how to use the determined abnormal parking event, please refer to Figure 5 , Figure 5 A flowchart of a method for determining an abnormal parking event based on a motion trajectory and adjusting a large model according to the embodiment of the disclosure, the flowchart 500 includes the following steps:
[0094] Step 501: screening motion trajectories with standing points from all motion trajectories to obtain screened motion trajectories;
[0095] This step provides a multi-condition composite judgment mechanism. First, a potential standing point is identified through speed-position joint analysis, for example, requiring the target speed to be below a threshold (such as 5 km / h) for N consecutive frames (usually N≥5, corresponding to 0.2 seconds / frame of monitoring video), and at the same time, the stop position can also be verified in combination with the high-precision map to determine whether it is in a no-parking area (excluding legal stops such as toll station queues). The above execution subject can establish a trajectory credibility evaluation model to perform kinematic backtracking verification on trajectories with temporary loss and reappearance, ensuring that the screening result is not disturbed by tracking interruption. For the unique traffic characteristics of the highway, a following distance analysis can also be introduced to avoid misjudging congestion slow-down as an abnormal parking.
[0096] Step 502: determining the standing time of the standing points existing in the screened motion trajectories;
[0097] This step aims to implement dynamic time window measurement according to the duration of stay, from the first time the trajectory point is below the speed threshold to the time when it accelerates again and exceeds the threshold, during which cubic spline interpolation can be used to compensate for possible detection gaps. To improve measurement accuracy, the above executing body can also refer to the surrounding vehicle motion state for calibration: when the vehicles in the adjacent lanes are all passing at high speed, the confirmation time for suspected stay is shortened, while in the overall slow-down section the judgment threshold is automatically extended. The target behavior characteristics during the stop period (such as the status of double flashing lights, car door opening and closing signals) are recorded as auxiliary basis for time length evaluation.
[0098] Step 503: generating a corresponding abnormal parking event for the corresponding traffic element according to the screened motion trajectory with a stay duration exceeding a preset duration;
[0099] This step aims to determine an abnormal parking event based on the screened motion trajectory with a stay duration exceeding a preset duration. Specifically, a differentiated event label can be generated according to the stay duration and the position risk level: for example, a short stay (10-30 seconds) is marked as a level three warning, triggering only local log recording; a long stay (30-120 seconds) generates a level two warning, notifying the road section management center; an ultra-long stay (more than 120 seconds) in high-risk areas such as curves and tunnels triggers a level one real-time warning, and the variable information board is linked to release avoidance information. The judgment process can integrate historical data to establish a dynamic baseline, for example, an average of 3 breakdowns per month occurs on a certain road section, and when the average is exceeded, the detection sensitivity is automatically increased.
[0100] Step 504: generating an abnormal parking video sample containing the abnormal parking event;
[0101] This step can use intelligent video editing technology to generate the abnormal parking video sample, for example, expanding 30 seconds forward and backward around the abnormal parking event to form a video clip, while embedding multiple layers of labeling information: the basic layer contains machine-readable data such as target position and speed curve; the semantic layer adds event labels such as "breakdown" and "illegal parking"; the environment layer records scene context such as weather and lighting. And the sample library construction can follow the principle of difficult sample priority, weight the abnormal events in low visibility scenes such as rain and fog, and ensure the balance of positive and negative samples (abnormal parking and normal driving).
[0102] Step 505: fine-tuning training a preset large autonomous driving model based on the abnormal parking video sample, to obtain an autonomous driving large model for providing autonomous driving services for autonomous driving vehicles in highway scenarios.
[0103] On the basis of step 504, the present step aims to fine-tune the automatic driving large model using the abnormal parking video sample. For example, the perception layer parameters of the automatic driving large model can be first frozen, only the risk assessment sub-network in the decision planning module is fine-tuned, the avoidance strategy under the sudden situation is simulated using the abnormal parking video sample, after preliminary stabilization, part of the visual backbone network is gradually unfrozen, and the detection ability of the abnormal parking target is enhanced. Further, in the training process, the adversarial sample enhancement technology can be used to improve the model robustness by simulating the parking scene under extreme weather.
[0104] The present embodiment provides a specific abnormal parking event determination scheme through steps 501-505, and a scheme of fine-tuning the training of the automatic driving large model using the abnormal parking event to enhance the automatic driving service.
[0105] On the basis of the above-mentioned embodiments, in addition to training the automatic driving large model using the determined abnormal parking event, the target automatic driving vehicle covering the area where the abnormal parking event is located in the planned driving route can be issued with an avoidance reminder, and then the accuracy of the determination of the abnormal parking event is verified according to the vehicle driving information of the target automatic driving vehicle when passing through the area, and the size of the confidence degree is corrected according to the determination accuracy.
[0106] That is, the abnormal event verification and confidence degree correction mechanism based on car-road cooperation provided by the present embodiment constitutes a dynamic optimization closed-loop feedback link. Among them, the intelligent issuance of the avoidance reminder can adopt a hierarchical push strategy to generate differentiated warning content according to the relative position and driving state (such as current speed, lane) of the target automatic driving vehicle and the abnormal event area. For vehicles more than 2 kilometers away from the event point, a complete data packet containing the event type, accurate coordinates and recommended lane change strategy is sent through the vehicle networking V2X channel; for vehicles that have entered the range of 1 kilometer, it is compressed into key vector information (such as "left front emergency lane fault vehicle, suggest to change lane to the right"), and transmitted through the 5G short delay channel. The above execution subject can predict the possible reaction path of the vehicle after receiving the information, when detecting that multiple vehicles receive the warning at the same time, automatically staggered sending and coordinating the lane change sequence to avoid secondary accidents caused by collective avoidance.
[0107] And the verification collection of driving information can be realized by constructing a multi-modal data fusion framework, receiving the perception data uploaded by the target vehicle when passing through the event area (such as laser radar point cloud, visual detection box), decision log (avoidance trajectory planning basis) and actual control signal (steering wheel angle, brake intensity). These information is spatio-temporally aligned with the original video recorded by the roadside monitoring system, and the existence state and spatial position difference of the abnormal parking target are verified through three-dimensional scene reconstruction technology. Special attention is paid to the data consistency of the automatic driving vehicle sensor and the roadside equipment, for example, when the vehicle-mounted camera does not detect the obstacle at the warning position, but the laser radar shows weak reflection points, it may indicate that there is a false alarm caused by a semi-transparent object (such as plastic cloth).
[0108] Finally, the dynamic correction of confidence can use the Bayesian probability updating model to convert the verification results (confirmation of existence, partial existence, false alarm) fed back by the vehicle into confidence adjustment factors. When multiple independent autonomous vehicles confirm the abnormal target, the event confidence is increased exponentially; and for contradictory feedback (such as A car reports confirmation, B car reports false alarm), the conflict analysis module is started, and after checking the influence factors such as sensor calibration state and environmental visibility, weighted correction is implemented. The corrected confidence not only affects the warning level of the current event, but also feeds back to the training link of the detection model mentioned above as the labeling basis for difficult samples. The above execution subject can maintain a confidence benchmark library based on road segment characteristics, set a higher verification passing threshold for areas prone to false alarms (such as dense reflective signs), and realize spatial adaptive precision control.
[0109] The implementation manner provided by the embodiment forms a continuously self-improving ecological system in actual operation: the roadside detection system provides macro event warning, the vehicle end feeds back micro verification data, and both sides realize co-evolution through the common language of confidence. When the detection system detects that a certain type of false alarm mode repeatedly occurs (such as continuously misjudging temporary construction signs as parked vehicles), it will automatically trigger targeted model retraining, and after updating, it will be verified through shadow mode to ensure the effectiveness of the correction measures. This two-way verification system of vehicle and road significantly improves the practical value of abnormal parking detection, enabling autonomous vehicles to deal with various sudden conditions on the highway in a progressive optimization manner.
[0110] For a better understanding, the disclosure also provides an obstacle level abnormal parking event detection scheme based on global complex scene understanding of time sequence video. The overall implementation process of the scheme can be seen from Figure 6a :
[0111] The whole framework mainly consists of multi-scale feature extraction module, spatial traffic element positioning and recognition module, relationship logic reasoning module, global scene understanding module and result prediction module. Among them:
[0112] 1) The function of the multi-scale feature extraction module is mainly to obtain the features of obstacles of different scales, providing sufficient guarantee for subsequent obstacle detection;
[0113] 2) The traffic element positioning and identification is mainly to provide position and category information for the foreground targets in the image;
[0114] 3) The relationship logic reasoning module is mainly to associate the same targets between different frames according to the category and position information;
[0115] 4) The global scene understanding module is mainly used to identify whether there is extreme weather such as rain, snow, and fog in the image, and whether there is a problem with the imaging quality of the image, such as screen flicker and ghosting;
[0116] 5) The result prediction module mainly gives a conclusion on whether there is an abnormal parking event in the video.
[0117] The inference process of the whole scheme is as follows:
[0118] Step one: obtain the video stream under the high-speed scene, and do simple data preprocessing operation;
[0119] Step two: send the processed video data into the multi-scale feature extraction module to obtain the multi-scale features of each frame of image;
[0120] Step three: send the multi-scale features of each frame of image into the global scene understanding module to obtain the meteorological attribute and image quality attribute prediction results; at the same time, send the results together with the original multi-scale features into the traffic element positioning and identification module to obtain the position, category information, meteorology, and image quality attribute of the traffic elements in each frame of image;
[0121] Step four: send the prediction results of each frame of image in step three into the relationship logic reasoning module to supplement the time sequence information of each obstacle;
[0122] Step five: according to the prediction information of step four, send into the result prediction module, output the abnormal parking event result.
[0123] In order to reflect the effect of the scheme provided in the embodiment, Figure 6b and Figure 6c is a group of effect comparison charts of whether to apply the abnormal parking detection method based on large model under the high-speed road scene provided in the present application provided by the embodiment of the present disclosure, and Figure 6d and Figure 6e is another group of effect comparison charts of whether to apply the abnormal parking detection method based on large model under the high-speed road scene provided in the present application provided by the embodiment of the present disclosure. Among them, Figure 6b and Figure 6dFor the abnormal event detection results of the scenarios not using the scheme provided by the present embodiment and using the conventional scheme, it can be seen that Figure 6b Abnormal parking events are still identified in the presence of screen flashing, Figure 6d In the case of overexposure, the ghosting is misidentified as an abnormal parking, and vice versa, after the scheme provided by the present embodiment is applied, Figure 6c The false identification of abnormal events is correctly eliminated due to the correct identification of the presence of screen flashing problems, Figure 6e The false identification of abnormal events is also correctly eliminated due to the correct identification of the ghosting problem.
[0124] Further referring to Figure 7 , as an implementation of the method shown in the above figures, the present disclosure provides an embodiment of an abnormal parking detection device based on a large model in a highway scene. The device embodiment corresponds to the method embodiment shown in Figure 2 , and the device can be applied to various electronic devices.
[0125] As shown in Figure 7 , the abnormal parking detection device based on a large model in a highway scene 700 of the present embodiment can include a road time sequence video acquisition unit 701, a multi-scale feature extraction unit 702, an image quality and environment information determination unit 703, a traffic element recognition unit 704, a motion trajectory determination unit 705, and an abnormal parking event determination unit 706. The road time sequence video acquisition unit 701 is configured to acquire a road time sequence video taken in a highway scene. The multi-scale feature extraction unit 702 is configured to perform a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain multi-scale features of each video frame. The image quality and environment information determination unit 703 is configured to determine image quality information and environment information of each video frame based on the multi-scale features of each video frame, respectively. The traffic element recognition unit 704 is configured to determine position information and category information of traffic elements present in each video frame based on the image quality information, the environment information, and the multi-scale features of each video frame, respectively. The image quality information and the environment information are used to determine the confidence of the position and category of the traffic elements. The motion trajectory determination unit 705 is configured to determine the motion trajectory of the same traffic elements based on the position information, the category information, and the time sequence information between the video frames. The abnormal parking event determination unit 706 is configured to determine the abnormal parking events present in the road time sequence video based on the motion trajectory of the traffic elements.
[0126] In the present embodiment, in the abnormal parking detection device 700 in the expressway scene based on a large model: the specific processing of the road time sequence video acquisition unit 701, the multi-scale feature extraction unit 702, the image quality and environment information determination unit 703, the traffic element identification unit 704, the motion trajectory determination unit 705, and the abnormal parking event determination unit 706 and the technical effects brought by the same can be referred to Figure 2 The related description of steps 201-206 in the corresponding embodiment will not be repeated here.
[0127] In some optional implementation manners of the present embodiment, in some other optional implementation manners of the present embodiment, the image quality and environment information determination unit 703 can include:
[0128] a feature fusion sub-unit configured to perform feature fusion processing on the multi-scale features of each video frame respectively, to obtain multi-scale fusion features of each video frame;
[0129] an initialization parameter sub-unit configured to initialize the learnable quality parameters and the learnable environment parameters corresponding to the respective video frames to have the same size and dimension number as the multi-scale fusion features according to the dimension number of the multi-scale fusion features;
[0130] a cross-attention sub-unit configured to perform cross-attention processing on the multi-scale fusion features of the same video frame and the learnable quality parameters and the learnable environment parameters, to obtain learned quality parameters and learned environment parameters;
[0131] a self-attention sub-unit configured to perform self-attention processing on the learned quality parameters and the learned environment parameters, to obtain optimized quality parameters and optimized environment parameters;
[0132] a multi-layer perception sub-unit configured to sequentially perform nonlinear transformation processing and classification prediction processing on the optimized quality parameters and the optimized environment parameters of each video frame respectively, to obtain image quality information and environment information of the respective video frame.
[0133] In some other optional implementation manners of the present embodiment, the feature fusion sub-unit can be further configured to:
[0134] for each video frame, the following feature fusion processing operation is performed:
[0135] perform average pooling processing on the original features of each scale of the video frame, to obtain length-unified features of each scale;
[0136] perform splicing processing on the length-unified features of each scale of the video frame, and determine the features obtained after splicing as the multi-scale fusion features of the respective video frame.
[0137] In some other optional implementations of the embodiment, the cross-attention subunit can be further configured to:
[0138] The cross-attention calculation is performed in the manner that the learnable quality parameter of the same video frame is taken as the query information, and the multi-scale fusion feature is taken as the key information and the value information, to obtain the learned quality parameter that learns the image quality related information from the multi-scale fusion feature;
[0139] The cross-attention calculation is performed in the manner that the learnable environment parameter of the same video frame is taken as the query information, and the multi-scale fusion feature is taken as the key information and the value information, to obtain the learned environment parameter that learns the environment related information from the multi-scale fusion feature.
[0140] In some other optional implementations of the embodiment, the self-attention subunit can be further configured to:
[0141] The first target feature with the highest correlation degree with the image quality in the learned quality parameter is determined, the correlation degree within the features is determined in the manner that the learned quality parameter is subjected to the self-attention processing, and other features with an actual correlation degree less than a first preset correlation degree threshold value from the first target feature are removed, to obtain the optimized quality parameter;
[0142] The second target feature with the highest correlation degree with the environment in the learned environment parameter is determined, the correlation degree within the features is determined in the manner that the learned environment parameter is subjected to the self-attention processing, and other features with an actual correlation degree less than a second preset correlation degree threshold value from the second target feature are removed, to obtain the optimized environment parameter.
[0143] In some other optional implementations of the embodiment, the multi-layer perception subunit can be further configured to:
[0144] For each video frame, the following multi-layer perception processing is performed:
[0145] The optimized quality parameter and the optimized environment parameter of the video frame are sequentially subjected to the processing of the first fully connected layer and the nonlinear transformation layer, respectively, to obtain the processed quality feature and the processed environment feature; wherein the first fully connected layer is used to increase the feature dimension of the optimized quality parameter and the optimized environment parameter;
[0146] The processed quality feature and the processed environment feature of the video frame are sequentially subjected to the processing of the second fully connected layer and the classification prediction layer, respectively, to obtain the image quality information containing the image quality category and the probability size, and the environment information containing the environment category and the probability size; wherein the second fully connected layer is used to reduce the feature dimension of the processed quality feature and the processed environment feature.
[0147] In some other optional implementations of the embodiment, the traffic element identification unit 704 can be further configured to:
[0148] determine the confidence of each video frame according to the image quality information and the environment information of the video frame, respectively;
[0149] identify the position information and the category information of the existing traffic element only from the video frames with the confidence exceeding the preset confidence threshold.
[0150] In some other optional implementations of the embodiment, the environment information includes weather sub-information and illumination intensity sub-information, the weather sub-information includes at least one of normal, rain, wind, fog, haze, thunder and lightning, and snow, and the illumination intensity sub-information includes at least one of normal, too bright, and too dark.
[0151] In some other optional implementations of the embodiment, the image quality information includes at least one of normal, screen flower, trailing, ghosting, overexposure, overdarkness, resolution lower than a preset resolution, and noise quantity exceeding a preset quantity.
[0152] In some other optional implementations of the embodiment, the road time sequence video acquisition unit 701 can be further configured to:
[0153] in response to the target highway scene corresponding to multiple road videos with different perspectives, performing perspective fusion on the multiple road videos to obtain the road time sequence video corresponding to the target highway scene after perspective completion.
[0154] In some other optional implementations of the embodiment, the abnormal parking event determination unit 706 can be further configured to:
[0155] screen the motion trajectories with the stay points from all the motion trajectories to obtain screened motion trajectories;
[0156] determine the stay duration of the stay points existing in the screened motion trajectories;
[0157] generate a corresponding abnormal parking event for a corresponding traffic element according to the screened motion trajectory with the stay duration exceeding the preset duration.
[0158] In some other optional implementations of the embodiment, the abnormal parking detection device 700 based on the large model in the highway scene can further include:
[0159] a sample generation unit configured to generate an abnormal parking video sample containing an abnormal parking event;
[0160] The fine-tuning training unit is configured to perform fine-tuning training on the preset large automatic driving model based on the abnormal parking video sample, to obtain the large automatic driving model for providing automatic driving services for the automatic driving vehicle in the highway scene.
[0161] In some other optional implementations of the present embodiment, the abnormal parking detection device 700 based on a large model in a highway scene can further include:
[0162] The avoidance reminding issuing unit is configured to issue an avoidance reminding to the target automatic driving vehicle covering the area where the abnormal parking event occurs in the planned driving route;
[0163] The determination accuracy verification unit is configured to verify the determination accuracy of the abnormal parking event according to the received vehicle driving information of the target automatic driving vehicle when passing through the area;
[0164] The confidence parameter correction unit is configured to correct the size of the confidence according to the determination accuracy.
[0165] The present embodiment exists as a device embodiment corresponding to the above-mentioned method embodiment. The abnormal parking detection device based on a large model in a highway scene provided by the present embodiment can first improve the recognition ability of traffic elements of different sizes in the video frame through the multi-scale feature extraction operation, and then determine the image quality information and environmental information of the corresponding video frame according to the multi-scale features, and then determine whether the traffic elements reflected by the multi-scale features exist, and the position and category of the traffic elements, so as to avoid the influence of image quality problems or environmental factors on the recognition accuracy of the traffic elements, so that the motion trajectory based on the traffic elements recognized according to the traffic elements can be determined in the complex scene whether there is an abnormal parking event in the highway scene, and then the corresponding avoidance is performed in time to ensure the driving safety of the object in the highway scene.
[0166] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, which comprises at least one processor and a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the abnormal parking detection method based on a large model in a highway scene described in any of the above embodiments.
[0167] According to the embodiments of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions for enabling a computer to implement the abnormal parking detection method based on a large model in a highway scene described in any of the above embodiments.
[0168] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which, when executed by a processor, can implement the large model-based abnormal parking detection method in a highway scene described in any of the above embodiments.
[0169] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0170] As shown in Figure 8 The electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0171] Various components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, and the like; an output unit 807, such as various types of displays, speakers, and the like; the storage unit 808, such as a magnetic disk, an optical disk, and the like; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0172] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 801 performs various methods and processes described above, such as the large model based abnormal parking detection method in highway scenarios. For example, in some embodiments, the large model based abnormal parking detection method in highway scenarios can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the large model based abnormal parking detection method in highway scenarios described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the large model based abnormal parking detection method in highway scenarios by any other appropriate means, such as by means of firmware.
[0173] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0174] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.
[0175] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0176] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0177] The systems and techniques described here can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0178] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS, Virtual Private Server) services.
[0179] According to the technical scheme of the embodiment of the present disclosure, first, the recognition ability of traffic elements of different sizes in the video frame is improved through the multi-scale feature extraction operation, and then the image quality information and the environment information of the corresponding video frame are determined according to the multi-scale features, and then the image quality information and the environment information of the video frame are used to determine whether the traffic elements reflected by the multi-scale features exist, and the position and category of the traffic elements, so as to avoid the influence of image quality problems or environmental factors on the recognition accuracy of the traffic elements, so that the motion trajectory based on the traffic elements identified according to this can determine whether there is an abnormal parking event in the highway scene even in a complex scene, and then the corresponding avoidance is performed in time to ensure the driving safety of the object in the highway scene.
[0180] It should be understood that various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical scheme of the present disclosure can be achieved, and the present disclosure is not limited herein.
[0181] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.
Claims
1. A method for detecting abnormal parking in a highway scene based on a large model, characterized in that, The method comprises the following steps: acquiring a road time sequence video captured by a highway scene; performing a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain a multi-scale feature of each video frame; performing feature fusion processing on the multi-scale feature of each video frame respectively to obtain a multi-scale fusion feature of each video frame; initializing a learnable quality parameter and a learnable environment parameter corresponding to the respective video frame with a dimension number consistent with the dimension number of the multi-scale fusion feature; performing cross-attention processing on the learnable quality parameter and the learnable environment parameter and the multi-scale fusion feature of the same video frame to obtain a learned quality parameter and a learned environment parameter; performing self-attention processing on the learned quality parameter and the learned environment parameter to obtain an optimized quality parameter and an optimized environment parameter; performing nonlinear transformation processing and classification prediction processing on the optimized quality parameter and the optimized environment parameter of each video frame in sequence to obtain image quality information and environment information of the respective video frame; determining the position information and the category information of the traffic elements existing in the respective video frame according to the image quality information, the environment information and the multi-scale feature of each video frame, wherein the image quality information and the environment information are used to determine the confidence of the position and the category of the traffic elements; determining the motion trajectory of the same traffic element according to the position information, the category information and the time sequence information between the video frames of the traffic elements existing in each video frame; determining an abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic elements.
2. The method of claim 1, wherein, The feature fusion processing on the multi-scale feature of each video frame to obtain the multi-scale fusion feature of each video frame comprises: for each video frame, the following feature fusion processing operation is performed: performing average pooling processing on the original features of each scale of the video frame to obtain length-unified features of each scale; performing splicing processing on the length-unified features of each scale of the video frame, and determining the feature obtained after splicing as the multi-scale fusion feature of the respective video frame.
3. The method of claim 1, wherein, The cross-attention processing of the learnable quality parameter and the learnable environment parameter and the multi-scale fusion feature of the same video frame to obtain the learned quality parameter and the learned environment parameter comprises: performing cross-attention calculation by taking the learnable quality parameter of the same video frame as query information and the multi-scale fusion feature as key information and value information to obtain the learned quality parameter learned from the multi-scale fusion feature and related to image quality; performing cross-attention calculation by taking the learnable environment parameter of the same video frame as query information and the multi-scale fusion feature as key information and value information to obtain the learned environment parameter learned from the multi-scale fusion feature and related to the environment.
4. The method of claim 1, wherein, The self-attention processing on the learned quality parameter and the learned environment parameter to obtain the optimized quality parameter and the optimized environment parameter comprises: determining a first target feature with the highest correlation degree with image quality in the learned quality parameter, determining the correlation degree within the features in a self-attention processing manner for the learned quality parameter, and eliminating other features with an actual correlation degree less than a first preset correlation degree threshold from the first target feature, to obtain the optimized quality parameter; determining a second target feature with the highest correlation degree with the environment in the learned environment parameter, determining the correlation degree within the features in a self-attention processing manner for the learned environment parameter, and eliminating other features with an actual correlation degree less than a second preset correlation degree threshold from the second target feature, to obtain the optimized environment parameter.
5. The method of claim 1, wherein, the optimized quality parameter and the optimized environment parameter of each video frame are sequentially subjected to nonlinear transformation processing and classification prediction processing to obtain image quality information and environment information of the corresponding video frame, including: for each video frame, the following multi-layer perception processing is performed: the optimized quality parameter and the optimized environment parameter of the video frame are sequentially subjected to first full connection layer and nonlinear transformation layer processing to obtain processed quality features and processed environment features; wherein the first full connection layer is used to increase the feature dimension of the optimized quality parameter and the optimized environment parameter; the processed quality features and the processed environment features of the video frame are sequentially subjected to second full connection layer and classification prediction layer processing to obtain image quality information containing image quality categories and probability sizes and environment information containing environment categories and probability sizes; wherein the second full connection layer is used to reduce the feature dimension of the processed quality features and the processed environment features.
6. The method of claim 1, wherein, the position information and the category information of the traffic elements existing in the corresponding video frame are determined according to the image quality information, the environment information and the multi-scale features of each video frame, the image quality information and the environment information are used to determine the confidence of the position and the category of the traffic elements, including: the confidence of each video frame is determined according to the image quality information and the environment information of the video frame; only the position information and the category information of the traffic elements existing in the video frame with a confidence exceeding a preset confidence threshold are identified.
7. The method according to any one of claims 1 to 6, characterized in that, the environment information includes weather sub-information and illumination intensity sub-information, the weather sub-information includes at least one of normal, rain, wind, fog, haze, lightning and snow, and the illumination intensity sub-information includes at least one of normal, too bright and too dark.
8. The method according to any one of claims 1 to 6, characterized in that, the image quality information includes at least one of normal, screen flower, trailing shadow, virtual shadow, overexposure, overdarkness, resolution lower than a preset resolution and noise quantity exceeding a preset quantity.
9. The method of claim 1, wherein, the road time sequence video shot in a highway scene is obtained, including: in response to the target highway scene corresponding to multiple road videos with different perspectives, the multiple road videos are subjected to perspective fusion to obtain a road time sequence video with perspective completion corresponding to the target highway scene.
10. The method of claim 1, wherein, The determining of the abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic element comprises: Screening the motion trajectories existing in the standing points to obtain screened motion trajectories; Determining the standing time of the standing points existing in the screened motion trajectories; According to the screened motion trajectories with the standing time exceeding the preset time length, generating the corresponding abnormal parking event for the corresponding traffic element.
11. The method of claim 10, wherein, Further comprising: Generating an abnormal parking video sample containing the abnormal parking event; Based on the abnormal parking video sample, fine-tuning training is performed on the preset large automatic driving model to obtain an automatic driving large model for providing automatic driving services for an automatic driving vehicle in a highway scene.
12. The method of claim 10, wherein, Further comprising: Issuing an avoidance reminder to a target automatic driving vehicle covering the area where the abnormal parking event is located in the planned driving route; According to the vehicle driving information of the target automatic driving vehicle received when the vehicle is driving through the area, verifying the determination accuracy of the abnormal parking event; According to the determination accuracy, correcting the size of the confidence.
13. An abnormal parking detection device in a highway scene based on a large model, characterized by, Comprise: A road time sequence video acquisition unit configured to acquire a road time sequence video shot in a highway scene; A multi-scale feature extraction unit configured to perform a multi-scale feature extraction operation on each video frame constituting the road time sequence video to obtain multi-scale features of each video frame; An image quality and environment information determination unit configured to perform feature fusion processing on the multi-scale features of each video frame respectively to obtain multi-scale fusion features of each video frame; initialize to obtain learnable quality parameters and learnable environment parameters corresponding to the corresponding video frame with the same dimension number as the dimension number of the multi-scale fusion features; cross attention processing is performed on the learnable quality parameters and the learnable environment parameters and the multi-scale fusion features of the same video frame to obtain learned quality parameters and learned environment parameters; self-attention processing is performed on the learned quality parameters and the learned environment parameters to obtain optimized quality parameters and optimized environment parameters; the optimized quality parameters and the optimized environment parameters of each video frame are sequentially subjected to nonlinear transformation processing and classification prediction processing to obtain image quality information and environment information of the corresponding video frame; A traffic element recognition unit configured to determine the position information and category information of the traffic element existing in each video frame according to the image quality information, the environment information and the multi-scale features of each video frame respectively, the image quality information and the environment information being used to determine the confidence of the position and the category of the traffic element; A motion trajectory determination unit configured to determine the motion trajectory of the same traffic element according to the position information, the category information and the time sequence information between the video frames of the traffic element existing in each video frame; An abnormal parking event determination unit configured to determine the abnormal parking event existing in the road time sequence video according to the motion trajectory of the traffic element.
14. An electronic device comprising: At least one processor; and A memory in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for detecting abnormal parking in expressway scene based on large model according to any one of claims 1-12. 15.A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method for detecting abnormal parking in expressway scene based on large model according to any one of claims 1-12. 16.A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method for detecting abnormal parking in expressway scene based on large model according to any one of claims 1-12.
Citation Information
Patent Citations
Image processing method and device, terminal and storage medium
CN110287778A
Highway vehicle abnormal behavior detection method based on deep learning
CN119991736A