Road condition prediction method and system based on computer vision technology
By building a practical topological network and using computer vision technology to predict driving road conditions in real time, the problems of large amount of calculation and inaccurate prediction in the existing technology are solved, and real-time and accurate prediction of road conditions are achieved.
Patent Information
- Application Number
- CN202510359584.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing road condition prediction methods include a large number of urban road networks, resulting in huge calculations and the inability to predict road conditions in real time and accurately.
Using a method based on computer vision technology, a practical topological network is built to predict driving road conditions in real time by obtaining the preset navigation route of the driving vehicle and the connected intersection.
Reduced the amount of calculation, real-time and accurate prediction of road conditions, and improved the driving experience.
Smart Images

Figure CN120148243A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of road condition prediction, and particularly to a road condition prediction method and system based on computer vision technology. Background Art
[0002] Road condition prediction is to predict the road conditions for the next period of time based on historical road condition information. At present, the ETA (Estimated Time of Arrival) deep model based on graph convolution for road condition prediction has been applied to navigation of Amap and Baidu, delivery of Meituan, pick-up and drop-off driving of Didi, etc.
[0003] Meanwhile, after retrieval, the latest prior art proposes a road condition prediction method, device, vehicle and storage medium (2024109162560). The method includes: obtaining broadcast information sent by a Bluetooth mesh network; extracting a network identifier and a target credential from the broadcast information, joining the Bluetooth mesh network based on the network identifier and the target credential, and determining a plurality of target other vehicle nodes in the Bluetooth mesh network according to the navigation route of the target vehicle; sending an active request to the plurality of target other vehicle nodes, wherein each target other vehicle node generates an assembled data packet according to the active request and transmits it to the target vehicle through the Bluetooth mesh network; receiving the assembled data packets sent by each target other vehicle node, and predicting the congestion condition in the navigation route according to the assembled data packets. Thus, the problems in the related art that in the case of complex roads or congested roads, the road conditions cannot be accurately predicted in a timely manner, resulting in problems such as the driving experience of users are solved.
[0004] In addition, after retrieval, the prior art provides a traffic road condition prediction method, a training method for a classification model and a computer-readable storage medium. The method includes: obtaining the cell load data and user terminal handover data of a target cell in a target time period; wherein the target cell is a cell that covers a target road section with signals; constructing a target feature vector according to the cell load data and the user terminal handover data; inputting the target feature vector into a preset classification model for identification to obtain the congestion level of the target road section in the target time period.
[0005] However, the technology of transmitting the above road condition prediction method, device, vehicle and storage medium to the target vehicle through the Bluetooth mesh network is limited by the transmission distance. In addition, the above traffic road condition prediction method, training method for a classification model and computer-readable storage medium are separated from the association between the driving vehicle and other vehicles, and cannot effectively give the driving road conditions corresponding to the preset navigation route of the driving vehicle. At the same time, the current road condition prediction method includes a large number of urban road networks, which results in a huge amount of calculation and cannot predict the road conditions in real time and accurately. Summary of the Invention
[0006] The present disclosure provides a technical solution for a road condition prediction method and system based on computer vision technology.
[0007] According to one aspect of the present disclosure, there is provided a road condition prediction method based on computer vision technology, including:
[0008] Obtain the preset navigation route corresponding to the preset starting point and preset ending point of the first preset driving vehicle, and respectively determine a plurality of intersection roads connected to the preset navigation route;
[0009] In real time, with the first preset driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, construct a real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads;
[0010] Based on the real-time variable topological network, predict the driving road conditions of the first preset driving vehicle corresponding to the preset navigation route.
[0011] Preferably, the step of constructing, in real time, with the first preset driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, a real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads includes:
[0012] In real time, with the first preset driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, determine the real-time variable topological range;
[0013] Within the real-time variable topological range, use the first preset driving vehicle and the intersection roads within the radius as nodes, and use the route corresponding to the intersection road closest to the first preset driving vehicle in the preset navigation route and the routes between the intersection roads within the radius as edges to construct a real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads.
[0014] Preferably, in the process of constructing, within the real-time variable topological range, a real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads with the first preset driving vehicle and the intersection roads within the radius as nodes and the route corresponding to the intersection road closest to the first preset driving vehicle in the preset navigation route and the routes between the intersection roads within the radius as edges, if there is a route in the direction of the first preset driving vehicle among the routes between the intersection roads within the radius, retain the edge corresponding to the real-time variable topological network; otherwise, delete the edge corresponding to the real-time variable topological network.
[0015] Preferably, before constructing the real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads, the method for determining the first preset driving vehicle corresponding to the preset navigation route in real time includes:
[0016] Obtain the multi - moment images corresponding to the video stream of the first set driving vehicle on the preset navigation route in real time;
[0017] Use a target detection model to detect the license plates of all driving vehicles in the multi - moment images, and determine the first set driving vehicle corresponding to the preset navigation route.
[0018] Preferably, the method of using a target detection model to detect the license plates of all set driving vehicles in the multi - moment images and determine the first set driving vehicle corresponding to the preset navigation route includes:
[0019] Use the target detection model to detect the license plates of all driving vehicles in the multi - moment images, and obtain a corresponding plurality of license plate detection frames;
[0020] Identify the license plate numbers corresponding to the inside of the plurality of license plate detection frames respectively, and obtain a corresponding plurality of recognized license plate numbers;
[0021] Based on the fact that the plurality of recognized license plate numbers are consistent with the preset license plate number of the first set driving vehicle corresponding to the preset navigation route, determine the driving vehicle corresponding to the consistent preset license plate number as the first set driving vehicle corresponding to the preset navigation route.
[0022] Preferably, if the use of the target detection model to detect the license plates of all driving vehicles in the multi - moment images fails to determine the first set driving vehicle corresponding to the preset license plate number of the preset navigation route, then respectively obtain the corresponding high - level multi - scale fusion features and low - level multi - scale fusion features for the multiple multi - scale features of the multi - moment images;
[0023] Perform grouped detail enhancement and fusion processing on the multiple multi - scale features, high - level multi - scale fusion features and low - level multi - scale fusion features respectively, and obtain a corresponding plurality of collaborative multi - scale fusion features;
[0024] Optimize the collaborative multi - scale fusion features respectively, and obtain a corresponding plurality of optimized collaborative multi - scale fusion features;
[0025] Detect or determine the first set driving vehicle in the multi - moment images respectively based on the plurality of optimized collaborative multi - scale fusion features.
[0026] Preferably, the method of predicting the driving road conditions corresponding to the preset navigation route based on the real - variable topological network includes:
[0027] Based on the real - variable topological network, determine a plurality of core nodes corresponding to the intersections in the driving direction on the preset navigation route;
[0028] Predict respectively whether multiple second preset vehicles in the real-time variable topological network will obstruct the first preset driving vehicle at the multiple core nodes;
[0029] If an obstruction occurs, predict the congestion condition of the corresponding intersection road in the real-time variable topological network based on the cumulative obstruction time of the multiple second preset vehicles and the preset time.
[0030] Preferably, the method for predicting respectively whether multiple second preset vehicles in the real-time variable topological network will obstruct the first preset driving vehicle at the multiple core nodes includes:
[0031] Calculate respectively multiple first distances and multiple second distances between the first preset driving vehicle and the multiple second preset vehicles in the real-time variable topological network and the multiple core nodes;
[0032] Predict multiple first speeds at which the first preset driving vehicle will reach the multiple core nodes corresponding to in the remaining driving distance;
[0033] Predict respectively multiple second speeds at which the multiple second preset vehicles will reach the multiple core nodes corresponding to;
[0034] Based on the multiple first distances and the multiple first speeds corresponding to the first preset driving vehicle, the multiple second distances corresponding to the multiple second preset vehicles and the multiple second speeds, determine whether each second preset vehicle in the multiple second preset vehicles will obstruct the first preset driving vehicle from passing through the corresponding core node.
[0035] Preferably, the method for determining whether each second preset vehicle in the multiple second preset vehicles will obstruct the first preset driving vehicle from passing through the corresponding core node based on the multiple first distances and the multiple first speeds corresponding to the first preset driving vehicle, the multiple second distances corresponding to the multiple second preset vehicles and the multiple second speeds includes:
[0036] Based on the multiple first distances and the multiple first speeds corresponding to the first preset driving vehicle, determine multiple first positions corresponding to the first preset driving vehicle;
[0037] Based on the multiple second distances corresponding to the multiple second preset vehicles and the multiple second speeds, determine multiple second positions corresponding to each second preset vehicle in the multiple second preset vehicles;
[0038] Based on multiple first positions corresponding to the first preset driving vehicle and multiple second positions corresponding to each of the multiple second preset vehicles, determine whether the multiple second preset vehicles obstruct the first preset driving vehicle at the multiple core nodes at the intersection corresponding to each core node.
[0039] Preferably, the method for determining whether the multiple second preset vehicles obstruct the first preset driving vehicle at the multiple core nodes at the intersection corresponding to each core node based on multiple first positions corresponding to the first preset driving vehicle and multiple second positions corresponding to each of the multiple second preset vehicles includes:
[0040] According to multiple first positions corresponding to the first preset driving vehicle and multiple second positions corresponding to each of the multiple second preset vehicles, determine whether each of the multiple second preset vehicles is respectively in front of the first preset driving vehicle on the driving road corresponding to the intersection corresponding to each core node; if so, the multiple second preset vehicles obstruct the first preset driving vehicle at the multiple core nodes; otherwise, the multiple second preset vehicles do not obstruct the first preset driving vehicle at the multiple core nodes.
[0041] Preferably, the method for predicting the congestion condition of the corresponding intersection in the real-time variable topology network based on the cumulative obstruction time of the multiple second preset vehicles includes:
[0042] Predict the obstruction time corresponding to each second preset vehicle, and determine the cumulative obstruction time according to the obstruction time;
[0043] If the cumulative obstruction time is less than the first preset time, predict that the corresponding intersection in the real-time variable topology network is not congested; otherwise, if the cumulative obstruction time is less than the second preset time, predict that the corresponding intersection in the real-time variable topology network is slightly congested; otherwise, predict that the corresponding intersection in the real-time variable topology network is severely congested.
[0044] Preferably, the method for predicting multiple first speeds at which the first preset driving vehicle reaches the multiple core nodes in the remaining driving distance includes:
[0045] Obtain the distance already traveled, the time already traveled, and the road condition corresponding to the remaining driving distance of the first preset driving vehicle on the preset navigation route; wherein, the road condition includes one or several of the number of first lanes corresponding to the driving direction of the first preset driving vehicle, the number of first vehicles corresponding to each lane, and the weather condition.
[0046] Based on the distance traveled, travel time, and road conditions corresponding to the remaining travel distance of the first-set driving vehicle on the preset navigation route, use a machine learning prediction model or a time series analysis algorithm to respectively predict the multiple first speeds corresponding to the first-set driving vehicle reaching the multiple core nodes during the remaining travel distance.
[0047] Preferably, the method for respectively predicting the multiple second speeds corresponding to the multiple second-set vehicles reaching the multiple core nodes includes:
[0048] Obtain the distance traveled, travel time, and road conditions corresponding to the remaining travel distance of the multiple second-set driving vehicles on the preset navigation route; wherein, the road conditions include one or several of the number of second lanes corresponding to the driving directions of the multiple second-set driving vehicles towards the multiple core nodes, the number of second vehicles corresponding to each lane, and the weather condition.
[0049] Based on the distance traveled, travel time, and road conditions corresponding to the remaining travel distance of the second-set driving vehicles on the preset navigation route, use a machine learning prediction model or a time series analysis algorithm to respectively predict the multiple second speeds corresponding to the first-set driving vehicle reaching the multiple core nodes during the remaining travel distance.
[0050] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including:
[0051] An acquisition unit, configured to acquire a preset navigation route corresponding to a preset starting point and a preset ending point of a first-set driving vehicle, and respectively determine a plurality of intersections connected to the preset navigation route;
[0052] A construction unit, configured to construct a real-time variable topological network corresponding to the intersections within the radius in real time with the first-set driving vehicle as the center and the real-time remaining travel distance corresponding to the preset navigation route as the radius;
[0053] A prediction unit, configured to predict the driving road conditions of the first-set driving vehicle on the preset navigation route based on the real-time variable topological network.
[0054] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: an electronic device, the electronic device is configured with a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to perform the above road condition prediction method.
[0055] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the above-mentioned road condition prediction method.
[0056] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: a computer program product, on which computer programs / instructions are provided, and when the computer programs / instructions are executed by a processor, the above-mentioned road condition prediction method is implemented.
[0057] In an embodiment of the present disclosure, a road condition prediction method and system based on computer vision technology are proposed to solve the technical problem that the current road condition prediction method includes a large number of urban road networks, which leads to a huge amount of calculation and cannot predict the road conditions in real time and accurately.
[0058] In an embodiment of the present disclosure, a new method for constructing a time-varying topological network is proposed. Among them, the construction method includes: a method of constructing a time-varying topological network corresponding to the intersection within the radius in real time with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, thereby alleviating the problem that the current road condition prediction method includes a large number of urban road networks, which leads to a huge amount of calculation, and solving the technical problem that the road conditions cannot be predicted in real time and accurately.
[0059] In addition, in the technical solution of the road condition prediction method and system based on computer vision technology, in an embodiment of the present disclosure, a new method for determining a driving vehicle is also proposed. Among them, the determination of the driving vehicle in this method includes: respectively obtaining corresponding high-level multi-scale fusion features and low-level multi-scale fusion features for multiple multi-scale features of multiple moment images corresponding to the video stream of the first set driving vehicle on the preset navigation route; respectively performing grouped detail enhancement and fusion processing on the multiple multi-scale features, high-level multi-scale fusion features and low-level multi-scale fusion features to obtain corresponding multiple collaborative multi-scale fusion features; respectively performing optimization processing on the collaborative multi-scale fusion features to obtain corresponding multiple optimized collaborative multi-scale fusion features; respectively detecting or determining the first set driving vehicle in the multiple moment images based on the multiple optimized collaborative multi-scale fusion features. The technical problem of difficult determination of the driving vehicle caused by multiple target driving vehicles, chaotic and complex backgrounds, target occlusion, and weak GPS signals is eliminated.
[0060] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and do not limit the present disclosure.
[0061] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings are incorporated herein and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0063] Figure 1 A flowchart showing a road condition prediction method based on computer vision technology according to an embodiment of the present disclosure;
[0064] Figure 2 A schematic diagram showing the construction principle of a real variable topology network according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0065] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0066] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior to or better than other embodiments.
[0067] The term "and / or" herein is merely a description of the associated relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0068] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0069] It can be understood that the above-mentioned various embodiments of the road condition prediction method based on computer vision technology mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further.
[0070] In addition, the present disclosure also provides a road condition prediction device or system, an electronic device, a computer-readable storage medium, and a program product based on computer vision technology, all of which can be used to implement any of the road condition prediction methods provided by the present disclosure. For the corresponding technical solutions and descriptions, refer to the corresponding records in the method section, which will not be elaborated here.
[0071] Figure 1 The flowchart showing the road condition prediction method based on computer vision technology according to an embodiment of the present disclosure is as follows Figure 1 As shown, the road condition prediction method based on computer vision technology includes: Step S101: Obtain a preset navigation route corresponding to a preset starting point and a preset ending point of a first preset driving vehicle, and respectively determine a plurality of intersection roads connected to the preset navigation route; Step S102: In real time, with the first preset driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, construct a real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads; Step S103: Based on the real-time variable topological network, predict the driving road conditions of the first preset driving vehicle on the preset navigation route. This can solve the technical problem that the current road condition prediction methods include a large number of urban road networks, resulting in a huge amount of calculation and inability to predict road conditions in real time and accurately.
[0072] Step S101: Obtain a preset navigation route corresponding to a preset starting point and a preset ending point of a first preset driving vehicle, and respectively determine a plurality of intersection roads connected to the preset navigation route.
[0073] In the embodiments of the present disclosure and other possible embodiments, the preset starting point and the preset ending point of the first preset driving vehicle include multiple navigation routes, and the optimal navigation route corresponding to the shortest time can be selected from the multiple navigation routes, and the optimal navigation route is configured as the preset navigation route.
[0074] Step S102: In real time, with the first preset driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, construct a real-time variable topological network corresponding to the intersection roads within the radius among the plurality of intersection roads.
[0075] In an embodiment of the present disclosure, centered on the first set driving vehicle in real time, with the real-time remaining driving distance corresponding to the preset navigation route as the radius, a real-time variable topological network corresponding to the intersections within the radius is constructed, including: Centering on the first set driving vehicle in real time, with the real-time remaining driving distance corresponding to the preset navigation route as the radius, determining the real-time variable topological range; within the real-time variable topological range, using the first set driving vehicle and the intersections within the radius as nodes, and using the route corresponding to the intersection closest to the first set driving vehicle on the preset navigation route and the routes between the intersections within the radius as edges, constructing a real-time variable topological network corresponding to the intersections within the radius among the multiple intersections.
[0076] Figure 2 Shows the construction schematic diagram of the real-time variable topological network according to an embodiment of the present disclosure. As Figure 2 shown, the preset navigation route corresponding to the preset starting point A and the preset ending point B of the first set driving vehicle C is configured as A-a1-a2-a3-B; where a1, a2, and a3 are respectively 3 intersections on the preset navigation route.
[0077] In an embodiment of the present disclosure and other possible embodiments, it further includes: optimizing the real-time variable topological network, and the optimization method includes: deleting the passed nodes of the first set driving vehicle in the real-time variable topological network and the nodes and the corresponding edges of the routes connected to the passed nodes and not on the preset navigation route.
[0078] In an embodiment of the present disclosure and other possible embodiments, during the driving process of the first set driving vehicle C, the real-time remaining driving distance corresponding to the first set driving vehicle C continuously decreases, and the real-time variable topological range corresponding to the real-time variable topological network also continuously decreases. For example, with the real-time remaining driving distance corresponding to the preset navigation route as the radius R, a real-time variable topological network corresponding to the intersections within the radius R among the multiple intersections is constructed. As Figure 2 shown, the nodes corresponding to the real-time variable topological network at this time are C, a2, a3, and c3. Therefore, the real-time variable topological network before optimization is C, a2, a3, and c3 and the corresponding edges of the routes between them. Optimizing the real-time variable topological network, deleting the passed node a2 of the first set driving vehicle in the real-time variable topological network and the nodes and the corresponding edges of the routes connected to the passed node a2 and not on the preset navigation route, obtaining the optimized real-time variable topological network.
[0079] In the embodiments of the present disclosure and other possible embodiments, within the variable topology range, taking the first preset driving vehicle and the intersections within the radius as nodes, and taking the route corresponding to the intersection closest to the first preset driving vehicle in the preset navigation route and the routes between the intersections within the radius as edges, during the process of constructing the variable topology network corresponding to the intersections within the radius among the multiple intersections, if there is a route in the direction of the first preset driving vehicle between the intersections within the radius corresponding to the route in the direction of the first preset driving vehicle, then retain the edge corresponding to the variable topology network; otherwise, delete the edge corresponding to the variable topology network.
[0080] In the embodiments of the present disclosure and other possible embodiments, when deleting the edge corresponding to the variable topology network, there is no route in the direction of the first preset driving vehicle between the intersections within the radius corresponding to the route in the direction of the first preset driving vehicle (there is no road from the intersection to the first preset driving vehicle, that is, a one-way street).
[0081] In the embodiments of the present disclosure and other possible embodiments, in summary, within the variable topology range, taking the first preset driving vehicle and the intersections within the radius as nodes, and taking the route corresponding to the intersection closest to the first preset driving vehicle in the preset navigation route and the routes between the intersections within the radius corresponding to the route in the direction of the first preset driving vehicle as edges, construct the variable topology network corresponding to the intersections within the radius among the multiple intersections.
[0082] In the embodiments of the present disclosure, before constructing the variable topology network corresponding to the intersections within the radius among the multiple intersections, the method for determining the first preset driving vehicle corresponding to the preset navigation route in real time includes: obtaining in real time the multi-moment images corresponding to the video stream of the first preset driving vehicle on the preset navigation route; using an object detection model to detect the license plates of all driving vehicles in the multi-moment images to determine the first preset driving vehicle corresponding to the preset navigation route.
[0083] In the embodiments of the present disclosure and other possible embodiments, obtaining the high-level multi-scale fusion feature and the low-level multi-scale fusion feature by using the multiple multi-scale features of the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route includes: performing multi-scale fusion processing on the set first N layers of features among the multiple multi-scale features to obtain the low-level multi-scale fusion feature; performing the multi-scale fusion processing on the set last N layers of features among the multiple multi-scale features to obtain the high-level multi-scale fusion feature.
[0084] In embodiments of the present disclosure and other possible embodiments, the performing of grouped detail enhancement and fusion processing on the multiple multi-scale features, high-level multi-scale fusion features, and low-level multi-scale fusion features to obtain multiple collaborative multi-scale fusion features includes: performing a first grouping strategy on the first n of the multiple multi-scale features, and performing the detail enhancement and fusion processing; performing a second grouping strategy on the multiple multi-scale features other than the n multiple multi-scale features and the low-level multi-scale fusion features, and performing the detail enhancement and fusion processing; performing a third grouping strategy on the high-level multi-scale fusion features, and performing the detail enhancement and fusion processing.
[0085] In embodiments of the present disclosure and other possible embodiments, the performing of a first grouping strategy on the first n of the multiple multi-scale features and performing the detail enhancement and fusion processing includes: performing detail enhancement processing on the i-th layer feature and the (i + 1)-th layer feature among the first n multiple multi-scale features to obtain a first detail enhancement feature; performing collaborative fusion processing on the i-th feature and the (i + 1)-th feature to obtain a first collaborative fusion feature; and obtaining a first collaborative multi-scale fusion feature corresponding to the first grouping based on the first collaborative fusion feature and the first detail enhancement feature.
[0086] In embodiments of the present disclosure and other possible embodiments, the performing of a second grouping strategy on the multiple multi-scale features other than the n multiple multi-scale features and the low-level multi-scale fusion features and performing the detail enhancement and fusion processing includes: performing collaborative fusion processing on the multiple multi-scale features of the multiple multi-scale features other than the n multiple multi-scale features and the low-level multi-scale fusion features to obtain a second collaborative fusion feature; performing detail enhancement processing on the low-level multi-scale fusion features to obtain a second detail enhancement feature; and obtaining a second collaborative multi-scale fusion feature corresponding to the second grouping based on the second collaborative fusion feature and the second detail enhancement feature.
[0087] In embodiments of the present disclosure and other possible embodiments, the performing of a third grouping strategy on the high-level multi-scale fusion features and performing the detail enhancement and fusion processing includes: performing collaborative fusion processing on the second collaborative multi-scale fusion feature corresponding to the second grouping and the high-level multi-scale fusion features to obtain a third collaborative fusion feature; performing detail enhancement processing on the high-level multi-scale fusion features to obtain a third detail enhancement feature; and obtaining a third collaborative multi-scale fusion feature corresponding to the second grouping based on the third collaborative fusion feature and the third detail enhancement feature.
[0088] In the embodiments of the present disclosure and other possible embodiments, the detail enhancement input feature of the detail enhancement process is configured as a first detail enhancement input feature, and the detail enhancement process includes: performing a first convolution operation on the first detail enhancement input feature to obtain a convolution input feature; performing multi-channel detail extraction feature extraction on the convolution input feature to obtain a plurality of detail extraction features; performing depth enhancement fusion processing on the plurality of detail extraction features and the convolution input feature to obtain a depth detail fusion feature; performing edge guidance processing on the depth detail fusion feature to obtain an edge detail enhancement feature.
[0089] In the embodiments of the present disclosure and other possible embodiments, the performing multi-channel detail extraction feature extraction on the convolution input feature to obtain the plurality of detail extraction features includes: performing three-channel detail enhancement processing on the convolution input feature respectively. The first-channel enhancement processing obtains a first sub-detail extraction feature and a second sub-detail extraction feature, the second-channel enhancement processing obtains a third sub-detail extraction feature and a fourth sub-detail extraction feature, and the third-channel obtains a fifth detail extraction feature; performing a multiplication operation on the first sub-detail extraction feature and the fourth sub-detail extraction feature to obtain a first detail extraction feature, and performing a multiplication operation on the second sub-detail extraction feature and the third sub-detail extraction feature to obtain a second detail extraction feature.
[0090] In the embodiments of the present disclosure and other possible embodiments, the performing depth enhancement fusion processing on the plurality of detail extraction features and the convolution input feature to obtain a depth detail fusion feature includes: performing dilated convolution and attention processing on a part of the plurality of detail extraction features to obtain corresponding first fusion sub-features; performing attention processing on another part of the plurality of detail extraction features to obtain corresponding second fusion sub-features; performing a connection operation on the first fusion sub-features and the convolution input feature to obtain a first connection feature; performing an addition operation on the first connection feature and the second fusion sub-features to obtain the depth detail fusion feature.
[0091] In the embodiments of the present disclosure and other possible embodiments, the performing edge guidance processing on the depth detail fusion feature to obtain an edge detail enhancement feature includes: performing an activation operation on the depth detail fusion feature to obtain a first depth fusion activation feature; performing multi-scale attention processing on the first depth fusion activation feature to obtain an attention feature; performing a residual convolution operation on the attention feature to obtain the edge detail enhancement feature, and the edge detail enhancement feature can be used as a first detail enhancement feature, a second detail enhancement feature or a third detail enhancement feature.
[0092] In embodiments of the present disclosure and other possible embodiments, the input of the collaborative fusion processing is configured as a second detail-enhanced input feature and a third detail-enhanced input feature. The collaborative fusion processing includes: performing attention processing on the second detail-enhanced input feature and the third detail-enhanced input feature respectively to obtain a first attention input feature and a second attention input feature; performing a multi-scale convolution operation on the first attention input feature to obtain a multi-scale convolution input feature; performing pooling processing on the second attention input feature, and performing the multi-scale convolution operation on the result of the pooling processing to obtain a multi-scale pooled convolution input feature; obtaining a collaborative multi-scale fusion feature corresponding to the second detail-enhanced input feature and the third detail-enhanced input feature based on the multi-scale convolution input feature and the multi-scale pooled convolution input feature.
[0093] In embodiments of the present disclosure and other possible embodiments, a camera or a camera can be used to obtain in real time multi-moment images corresponding to the video stream of the first set driving vehicle on the preset navigation route.
[0094] In embodiments of the present disclosure and other possible embodiments, the target detection model can be configured as a target detection model corresponding to YOLO. The target detection model corresponding to YOLO is used to detect the license plates of all driving vehicles in the multi-moment images to determine the first set driving vehicle corresponding to the preset navigation route.
[0095] In an embodiment of the present disclosure, the method for using the target detection model to detect the license plates of all set driving vehicles in the multi-moment images to determine the first set driving vehicle corresponding to the preset navigation route includes: using the target detection model to detect the license plates of all driving vehicles in the multi-moment images to obtain corresponding multiple license plate detection frames; respectively identifying the license plate numbers corresponding to the multiple license plate detection frames to obtain corresponding multiple identified license plate numbers; if the multiple identified license plate numbers are consistent with the preset license plate number of the first set driving vehicle corresponding to the preset navigation route, determining the driving vehicle corresponding to the consistent preset license plate number as the first set driving vehicle corresponding to the preset navigation route.
[0096] In an embodiment of the present disclosure, if the license plates of all driving vehicles in the multi-moment images are detected by using the target detection model and the first set driving vehicle corresponding to the preset license plate number consistent with the preset navigation route is not determined, the multi-scale features of the multi-moment images are respectively processed to obtain corresponding high-level multi-scale fusion features and low-level multi-scale fusion features; the grouping detail enhancement and fusion processing are respectively performed on the multi-scale features, high-level multi-scale fusion features and low-level multi-scale fusion features to obtain corresponding multiple collaborative multi-scale fusion features; the collaborative multi-scale fusion features are respectively optimized to obtain corresponding multiple optimized collaborative multi-scale fusion features; and the first set driving vehicle in the multi-moment images is respectively detected or determined based on the multiple optimized collaborative multi-scale fusion features.
[0097] In the embodiments of the present disclosure and other possible embodiments, after obtaining the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route, by performing multiple multi-scale feature extractions on the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route, the quantification of the information of different scale features can be realized, providing a basis for the expression of the features for subsequent detail extraction.
[0098] In the embodiments of the present disclosure and other possible embodiments, a specific feature extraction network can be used to perform multiple multi-scale feature extractions on each image in the multi-moment image group corresponding to the video stream of the driving vehicle on the preset navigation route. Among them, the feature extraction network can be a residual network, a pyramid network, a U-net network, or other network structures capable of performing feature extractions of different scales and different levels. The present disclosure does not make specific limitations on this. For different scale features, further low-level and high-level feature fusions at different scales can be performed to obtain fusion features at a single scale and a composite scale.
[0099] In the embodiments of the present disclosure and other possible embodiments, by performing feature fusion on feature maps with different resolutions, the recognition ability of the driving vehicle (the first set driving vehicle) on the preset navigation route can be enhanced. Specifically, the feature fusion module in the embodiments of the present disclosure can combine high-level semantic information, middle-level structural information, and low-level detail information, and effectively integrate multiple multi-scale features through hierarchical upsampling and compression operations, thereby generating a fusion feature map with rich representation ability, which can more accurately locate and identify the area of the driving vehicle (the first set driving vehicle) on the preset navigation route.
[0100] In the embodiments of the present disclosure and other possible embodiments, the fusion module can be optimized through a predefined training strategy to enable it to adapt to the diverse characteristics of a driving vehicle (the first set driving vehicle) on a preset navigation route in different scenarios. Alternatively, based on the scene category of the image group, the driving vehicle characteristics (the first set driving vehicle characteristics) on different categories of preset navigation routes can be classified, thereby improving the generalization ability of multiple multi-scale feature fusions. Or, in combination with the attention mechanism based on region extraction, the expression ability of the fused feature map for the driving vehicle area (the first set driving vehicle area) on the preset navigation route is further strengthened, enabling the model to more accurately capture the detailed extraction features of the driving vehicle (the first set driving vehicle characteristics) on the preset navigation route. These fused feature maps can be used as the input of the target detection module, and finally generate a target detection prediction map of the driving vehicle (the first set driving vehicle characteristics) on the preset navigation route, which is used to represent the area and characteristics of the driving vehicle (the first set driving vehicle characteristics) on the preset navigation route.
[0101] In the embodiments of the present disclosure and other possible embodiments, grouping detail enhancement and fusion processing are performed on the multiple multi-scale features, high-level multi-scale fusion features, and low-level multi-scale fusion features to obtain multiple collaborative multi-scale fusion features; after obtaining features of different scales, high-level multi-scale fusion features, and low-level multi-scale fusion features, each feature can be grouped, and detail enhancement and fusion processing are performed on the grouped features. Based on the understanding of common features, collaborative multi-scale fusion features of different scales can be extracted. Through the detail enhancement process, the corresponding detail enhancement features of different features can be obtained. Combining the channel attention mechanism and the spatial attention mechanism, further hierarchical feature fusion and edge processing are performed on the detail enhancement features to obtain more optimized collaborative multi-scale fusion features; this feature prediction map can provide fine features of the target area, effectively suppress background interference, and improve the detection accuracy of the driving vehicle on the preset navigation route.
[0102] In the embodiments of the present disclosure and other possible embodiments, optimization processing is respectively performed on the collaborative multi-scale fusion features to obtain multiple optimized collaborative multi-scale fusion features. The optimization processing includes, but is not limited to, using multi-scale convolution kernels to extract features of different sizes and directions in the image, fusing features of different scales based on the adaptive attention mechanism, and further optimizing the obtained feature information through the optimization processing to improve the detection ability of the driving vehicle on the preset navigation route.
[0103] In the embodiments of the present disclosure and other possible embodiments, based on the multiple optimized collaborative multi-scale fusion features, the driving vehicle in the multi-moment images corresponding to the video stream on the preset navigation route is detected. The obtained optimized features can be connected, combined, or fused to obtain information capable of expressing features extracted with different details, and finally an accurate positioning prediction map is obtained. The multi-level feature fusion, detail enhancement processing, and common feature optimization fusion are integrated into a unified framework for accurate detection. Through multiple multi-scale feature fusion operations, information at different scales can be captured; through detail enhancement processing, the detail expression in the feature map can be optimized and background interference can be suppressed; through common feature extraction and fusion processing, the global information and local details in the image group can be effectively combined. Finally, by using the method of joint optimization to integrate various feature information, an accurate positioning prediction map is achieved, significantly improving the accuracy and robustness of the detection.
[0104] In the embodiments of the present disclosure and other possible embodiments, it is used to perform fusion processing on the multi-level feature maps of the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route to obtain multiple multi-scale feature fusion maps; it is used to refine the local features to generate detail-enhanced features, and then perform collaborative fusion on the feature maps with different resolutions to obtain collaborative multi-scale fusion features; it is used to perform joint optimization on the fused feature maps, and finally generate a positioning prediction map including each image.
[0105] In the embodiments of the present disclosure and other possible embodiments, before performing object detection, the embodiments of the present disclosure first need to obtain the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route. The multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route are an image group including at least one image. Each image group may include at least one hidden object that is difficult to discover and detect. The object types in each image of each image group may be the same or different. When training and constructing the detection model corresponding to the detection method, the object types in each image group are the same, so as to facilitate the extraction of common features by each part of the structure that can accurately express the driving vehicle on the same preset navigation route.
[0106] In the embodiments of the present disclosure and other possible embodiments, in the case of obtaining multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route, the embodiments of the present disclosure may perform multiple multi-scale feature extractions and feature fusions at different levels and depths on the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route. Among them, the multiple multi-scale feature extractions in the embodiments of the present disclosure can be implemented by a feature extraction module. For example, it can be a feature extraction network, and the feature extraction network includes a backbone network for processing the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route. The backbone network of the embodiments of the present disclosure can be implemented by using feature extraction networks such as a residual network and a pyramid network. The backbone network is used to perform feature extraction on the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route to obtain multiple multi-scale features at high and low levels.
[0107] In the embodiments of the present disclosure and other possible embodiments, after obtaining features at different scales, the features can be fused to obtain fused features. The embodiments of the present disclosure can use multiple multi-scale features of the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route to obtain high-level multi-scale fused features and low-level multi-scale fused features. The multiple multi-scale features include image features at at least three scales. Specifically, the embodiments of the present disclosure can fuse the low-level features in a preset manner to obtain low-level multi-scale fused features, and fuse the high-level features to obtain high-level multi-scale fused features. In the embodiments of the present disclosure, the use of multiple multi-scale features of the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route to obtain high-level multi-scale fused features and low-level multi-scale fused features includes: performing multi-scale fusion processing on the set first N layers of features among the multiple multi-scale features to obtain the low-level multi-scale fused features; performing the multi-scale fusion processing on the set last N layers of features among the multiple multi-scale features to obtain the high-level multi-scale fused features; N is an integer greater than or equal to 2. In the embodiments of the present disclosure, there is an intersection allowed between the first N layers and the set last N layers of features. For example, when there are 3 multiple multi-scale features and N is 2, the first 2 layers of multiple multi-scale features form low-level multi-scale fused features, and the last 2 layers of multiple multi-scale features form high-level multi-scale fused features. The second layer of multiple multi-scale features is used to constitute both low-level multi-scale fused features and high-level multi-scale fused features at the same time.
[0108] In the embodiments of the present disclosure and other possible embodiments, multiple methods are used to fuse and generate new high-level multi-scale fusion features and new low-level multi-scale fusion features. For example, at least two methods can be used to form high-level multi-scale fusion features and low-level multi-scale fusion features respectively, and new high-level multi-scale fusion features are generated based on the convolution results of at least two high-level multi-scale fusion features, and new low-level multi-scale fusion features are generated based on the convolution results of at least two low-level multi-scale fusion features, which can further ensure the integrity, accuracy, and robustness of feature information. N can be configured as 2 and 3 to obtain two high-level multi-scale fusion features and two low-level multi-scale fusion features. Further, a convolution operation is performed on the two high-level multi-scale fusion features to obtain new high-level multi-scale fusion features, and a convolution operation is performed on the two low-level multi-scale fusion features to obtain new low-level multi-scale fusion features.
[0109] In the embodiments of the present disclosure and other possible embodiments, the input multi-level feature maps are fused to capture multi-scale information from low level to high level, and multiple multi-scale feature fusion maps are obtained. The feature fusion of multiple multi-scale features in the embodiments of the present disclosure may include: in the order of multiple multi-scale features from high to low, the high-level features are upsampled to have the same resolution as the adjacent low-level features, and the two are combined through channel splicing to obtain a fusion feature map of the two; the fusion feature is further upsampled to the resolution of the low-level features, and channel compression is performed through a convolutional layer to obtain a feature map with the same resolution as the adjacent low-level features and an optimized number of channels; the feature is spliced with the low-level features to obtain the corresponding fusion feature. Through the above method, the low-level multi-scale fusion features and high-level multi-scale fusion features after the fusion of the first N layers and the last N layers can be obtained.
[0110] In the embodiments of the present disclosure and other possible embodiments, the high-level features are upsampled to have the same resolution as the middle-level features, and then the two are combined through channel splicing. This step retains the global semantic information of the high-level features and the texture information of the middle-level features. Subsequently, the fused features are further upsampled to the resolution of the low-level features, and channel compression is performed through a 1×1 convolutional layer to ensure that the representation of the fused features is more compact. These fused features are then spliced with the low-level features to form the final output. This process gradually fuses multi-scale information, utilizes the global receptive field of the high-level features, the local patterns of the middle-level features, and the detail resolution ability of the low-level features, providing a more comprehensive and robust feature representation for detection.
[0111] In the embodiments of the present disclosure and other possible embodiments, in the case of obtaining multiple multi-scale features, low-level multi-scale fusion features, and high-level multi-scale fusion features, detail enhancement and collaborative fusion processing can be further performed. The feature extraction network includes four branches for extracting feature information at different levels. By performing grouped detail enhancement and collaborative fusion processing on the obtained features, the feature detail expression and collaborative fusion can be adaptively performed according to the feature characteristics, and the effective feature information in the multi-moment images corresponding to the video stream of the driving vehicle on the preset navigation route can be extracted to the greatest extent. The performing of grouped detail enhancement and fusion processing on the multiple multi-scale features, high-level multi-scale fusion features, and low-level multi-scale fusion features to obtain multiple collaborative multi-scale fusion features includes: performing a first grouping strategy on the first n multiple multi-scale features among the multiple multi-scale features, and performing the detail enhancement and fusion processing; performing a second grouping strategy on the multiple multi-scale features other than the n multiple multi-scale features and the low-level multi-scale fusion features, and performing the detail enhancement and fusion processing; performing a third grouping strategy on the high-level multi-scale fusion features, and performing the detail enhancement and fusion processing; where n is an integer greater than 1.
[0112] In the embodiments of the present disclosure and other possible embodiments, grouping can be performed according to the characteristics of the obtained feature information. For low-level features, in the embodiments of the present disclosure, at least two adjacent layer features can be used as a group for detail enhancement and collaborative fusion processing. For high-level features, in the embodiments of the present disclosure, the high-level features and the obtained low-level multi-scale fusion features can be used as a group for detail enhancement and collaborative fusion processing. For high-level multi-scale fusion features, since they contain middle-level and high-level information, they can be directly used for detail enhancement and collaborative fusion processing with the previously obtained middle-low-level features.
[0113] In embodiments of the present disclosure and other possible embodiments, multiple multi-scale features, low-level multi-scale fusion features, and high-level multi-scale fusion features can be divided into multiple groups. Among them, the i-th layer feature and the (i + 1)-th layer feature in the first n multiple multi-scale features can be determined as the first group, and the multiple multi-scale features other than the first n multiple multi-scale features and the low-level multi-scale fusion features can be determined as the second group; the high-level multi-scale fusion features and the output of the second group can be determined as the third group. In one example, the first group includes two cases. In the first case, the first group includes the first feature and the second feature. In the second case, the first group includes the second feature and the third feature. The second group includes the fourth feature and the low-level multi-scale fusion features, and the fourth group includes the high-level multi-scale fusion features. The first group mainly processes low-level features, which have high resolution and rich edge details. Directly completing the interaction through the collaborative fusion feature fusion module can meet the target detection requirements. The second group and the third group process deep features. The second group combines multiple multi-scale feature fusions and integrates medium and low-level feature information to enhance its ability to express spatial details; the high-level multi-scale fusion features of the third group can be fused with the processed features of the second group, thereby aggregating the features of all layers and paying more attention to the extraction of global semantic information. Based on the above, the embodiments of the present disclosure can respectively achieve effective fusion of low-level, medium-low-level, and all-scale information.
[0114] In embodiments of the present disclosure and other possible embodiments, a first grouping strategy is executed on the first n multiple multi-scale features among the multiple multi-scale features, and the detail enhancement and fusion processing is performed, including: determining the i-th layer feature and the (i + 1)-th layer feature in the first n multiple multi-scale features as the first group; performing detail enhancement processing on the i-th layer feature and the (i + 1)-th layer feature in the first n multiple multi-scale features to obtain a first detail enhancement feature; performing collaborative fusion processing on the i-th feature and the (i + 1)-th feature to obtain a first collaborative fusion feature; and obtaining a first collaborative multi-scale fusion feature corresponding to the first group based on the connection feature of the first collaborative fusion feature and the first detail enhancement feature. Among them, the detail enhancement processing can be implemented by a detail enhancement module, and the collaborative fusion processing can be implemented by a collaborative fusion feature fusion module.
[0115] In embodiments of the present disclosure and other possible embodiments, the i-th feature and the (i + 1)-th feature can be the first feature and the second feature respectively, and can also be the second feature and the third feature. Through the above embodiments, the first collaborative multi-scale fusion features in two cases can be obtained, so as to obtain the first collaborative multi-scale fusion features after detail enhancement and collaborative fusion at different scales.
[0116] In addition, in the embodiments of the present disclosure, performing a second grouping strategy on multiple multi-scale features other than the n multiple multi-scale features and the low-level multi-scale fusion features, and performing the detail enhancement and fusion processing includes: determining the multiple multi-scale features other than the n multiple multi-scale features and the low-level multi-scale fusion features as the second group; performing collaborative fusion processing on the multiple multi-scale features of the second group to obtain a second collaborative fusion feature; performing detail enhancement processing on the low-level multi-scale fusion features to obtain a second detail enhancement feature; and obtaining a second collaborative multi-scale fusion feature corresponding to the second group based on the connection feature of the second collaborative fusion feature and the second detail enhancement feature. In the embodiments of the present disclosure, the fourth feature and the low-level multi-scale fusion feature can be used as the second group, and through this embodiment, the collaborative fusion of high-level features and middle-low-level features can be realized.
[0117] In the embodiments of the present disclosure and other possible embodiments, performing a third grouping strategy on the high-level multi-scale fusion feature and performing the detail enhancement and fusion processing includes: determining the high-level multi-scale fusion feature and the output of the second group as the third group; performing collaborative fusion processing on the output of the second group and the high-level multi-scale fusion feature to obtain a third collaborative fusion feature; performing detail enhancement processing on the high-level multi-scale fusion feature to obtain a third detail enhancement feature; and obtaining a third collaborative multi-scale fusion feature corresponding to the second group based on the third collaborative fusion feature and the third detail enhancement feature.
[0118] In the embodiments of the present disclosure and other possible embodiments, the obtained high-level multi-scale fusion feature can be subjected to detail enhancement and further collaborative fusion processing. The input of the collaborative fusion feature fusion module of the third group is different from that of other groups. It connects two results output from the multiple multi-scale feature fusion modules (the output from the second group and the high-level fusion output). This design is because the third group aggregates the features of all layers and needs to fully integrate the detail information from the second group and its own global semantic information through the collaborative fusion feature fusion module to ensure that the final output can achieve a balance in detail and semantic expression, thereby more accurately capturing the global characteristics. This differential design effectively coordinates the feature complementarity between different branches and improves the overall performance of the model.
[0119] In the embodiments of the present disclosure and other possible embodiments, the features of each group are first enhanced by a detail enhancement module to strengthen the local features, then enter the collaborative fusion feature fusion module to interact with the features of other layers, and finally pass through the joint optimization module for joint output. The detail enhancement processing and collaborative fusion processing implemented in the present disclosure are described in detail below. The detail enhancement model can realize the ability to capture details, jointly downsample the edge information and fuse it with the deep features to improve the ability to extract edge information.
[0120] In the embodiments of the present disclosure and other possible embodiments, the detail enhancement input feature of the detail enhancement process is configured as a first detail enhancement input feature, and the detail enhancement process includes: performing a first convolution operation on the first detail enhancement input feature to obtain a convolution input feature; performing multi-channel detail extraction feature extraction on the convolution input feature to obtain a plurality of detail extraction features; performing depth enhancement fusion processing on the plurality of detail extraction features and the convolution input feature to obtain a depth detail fusion feature; and performing edge guidance processing on the depth detail fusion feature to obtain an edge detail enhancement feature.
[0121] In the embodiments of the present disclosure and other possible embodiments, the outputs of four branches formed by three groups can be configured as the first detail enhancement input feature, and the first detail enhancement input feature is transmitted to a detail enhancement module for detail enhancement processing, where feature and edge feature enhancement optimization can be performed to achieve accurate detection.
[0122] Among them, in the embodiments of the present disclosure and other possible embodiments, first, a first convolution operation can be performed on the first detail enhancement input feature to obtain a convolution input feature, and this first convolution operation can be implemented through a convolution layer (1*1 convolution, normalization configured after the convolution, and an activation function configured after the normalization). Further, multi-channel detail extraction feature extraction can be performed on the convolution input feature, and by combining the detail extraction feature results from multiple angles, the extraction ability of the detail extraction features can be enhanced.
[0123] In the embodiments of the present disclosure and other possible embodiments, the performing multi-channel detail extraction feature extraction on the convolution input feature to obtain the plurality of detail extraction features includes: respectively performing three-way detail enhancement processing on the convolution input feature, where the first-way enhancement processing obtains a first sub-detail extraction feature and a second sub-detail extraction feature, the second-way enhancement processing obtains a third sub-detail extraction feature and a fourth sub-detail extraction feature, and the third-way obtains a fifth detail extraction feature; performing a multiplication process on the first sub-detail extraction feature and the fourth sub-detail extraction feature to obtain a first detail extraction feature, and performing a multiplication process on the second sub-detail extraction feature and the third sub-detail extraction feature to obtain a second detail extraction feature.
[0124] In the embodiments of the present disclosure and other possible embodiments, the detail enhancement process includes three processes: multi-scale enhancement, feature enhancement and fusion, and edge guidance fusion. In the multiple multi-scale feature enhancement processes, corresponding detail extraction features can be obtained respectively starting from three-way feature processing. In terms of multiple multi-scale feature enhancement, through the combination of convolution kernels of multiple different scales, features of different sizes and directions in the image can be extracted, and this design enables the model to perceive significant features at multiple scales.
[0125] Among them, in the embodiments of the present disclosure and other possible embodiments, the first - path detail enhancement processing includes: first, performing adaptive max - pooling processing on the convolutional input features, then successively performing enhancement processing of detail - extraction features on the convolutional input features using convolutional kernels of 1×3, 3×5, and 5×1, and then performing a dilated convolution operation with a dilation rate of 3 and a convolutional kernel of 3×3 to obtain the first sub - detail - extraction feature. At the same time, successively performing enhancement processing of detail - extraction features on the convolutional input features using 1×5, 5×3, and 3×1, and then performing a dilated convolution operation with a dilation rate of 5 and a convolutional kernel of 3×3 to obtain the second sub - detail - extraction feature.
[0126] In the embodiments of the present disclosure and other possible embodiments, the second - path detail enhancement processing includes: first, performing adaptive max - pooling processing on the convolutional input features, then successively performing enhancement processing of detail - extraction features on the convolutional input features using convolutional kernels of 1×7, 7×3, and 3×1, and then performing a dilated convolution operation with a dilation rate of 5 and a convolutional kernel of 3×3 to obtain the third sub - detail - extraction feature. At the same time, successively performing enhancement processing of detail - extraction features on the convolutional input features using 1×7, 7×3, and 3×1, and then performing a dilated convolution operation with a dilation rate of 1 and a convolutional kernel of 3×3 to obtain the fourth sub - detail - extraction feature.
[0127] In the embodiments of the present disclosure and other possible embodiments, the third - path detail enhancement processing includes: performing a convolution operation with a 3×3 convolutional kernel on the convolutional input features to obtain the fifth detail - extraction feature.
[0128] Based on the above, in the embodiments of the present disclosure and other possible embodiments, a multi - branch convolutional structure can be adopted. By introducing a combination form of multi - scale convolutional kernels, convolution operations with different receptive fields are fused simultaneously. In addition, learning of asymmetric feature patterns is realized, and this design can better adapt to the irregularity of the target shape and the inconsistency of directions.
[0129] More specifically, there are 6 branches in total for the multiple multi-scale feature enhancement parts. First, a 1×1 convolution operation is used to reduce the number of channels by half. Then, after using adaptive maximum pooling kernels in the first and second branches formed by the first path, and the fourth and fifth branches formed by the second path, in the first branch, convolutional layers of 1×5, 5×3, 3×1 and a convolutional layer of 3×3 with a dilation rate of 3 are used; in the second branch, convolutional layers of 1×3, 3×5, 5×1 and a convolutional layer of 3×3 with a dilation rate of 5 are used; in the fourth branch, convolutional layers of 1×3, 3×7, 7×1 and a convolutional layer of 3×3 with a dilation rate of 7 are used; in the fifth branch, convolutional layers of 1×7, 7×3, 3×1 and a convolutional layer of 3×3 with a dilation rate of 3 are used; the third branch is a shortcut branch that directly outputs the convolutional input features; the sixth branch formed by the third path consists of a 3×3 convolution operation. Each branch has different convolutional sizes and pooling strategies, which can capture information at different scales, directions and positions in the image. Therefore, multiple branches can make the features of the network more diverse and expressive. The design concept of this module is to retain both local details and global semantic information.
[0130] In the embodiments of the present disclosure and other possible embodiments, when four sub-detail extraction features are obtained, cross-enhancement processing can be further performed. Specifically, a multiplication process is performed on the first sub-detail extraction feature and the fourth sub-detail extraction feature to obtain a first detail extraction feature, and a multiplication process is performed on the second sub-detail extraction feature and the third sub-detail extraction feature to obtain a second detail extraction feature. Similarly, a multiplication process is performed on the fourth sub-detail extraction feature and the first sub-detail extraction feature to obtain a third detail extraction feature, and a multiplication process is performed on the third sub-detail extraction feature and the second sub-detail extraction feature to obtain a fourth detail extraction feature.
[0131] In the embodiments of the present disclosure and other possible embodiments, in terms of feature enhancement and fusion, an adaptive attention mechanism strategy is proposed to integrate features of different scales into a unified representation. By introducing a multi-scale channel attention mechanism and a module that combines spatial and channel attention mechanisms, the distinctiveness of features is further enhanced in both the spatial and channel dimensions. This innovative design ensures that the model can dynamically allocate computational resources, focus on key regions, and effectively suppress the impact of background noise on the detection performance. More specifically, in the first, second, fourth, and fifth branches, 3×3 convolutional layers with different dilation rates and a module that combines spatial and channel attention mechanisms are added to the output results of this part. Finally, a concatenation operation is performed with the output results of the feature enhancement and fusion part in the third branch. The subsequent results then pass through a 1×1 convolutional kernel and are added to the sixth branch that only passes through the module that combines spatial and channel attention mechanisms. Finally, the output is obtained through an activation function and multi-scale channel attention. In this part, the features are concatenated and then further extracted through convolution to ensure the consistency of the feature space in the interaction method. The final interaction is responsible for fusing the features extracted from the above branches into a unified feature representation, which can complement the information of local details and global context, enhance the comprehensive understanding of the target area, and suppress irrelevant information, such as background noise or redundant local details.
[0132] In the embodiments of the present disclosure and other possible embodiments, multiple branches can balance the relationship between the two through different scale and convolutional kernel strategies, so that the refinement of local features does not sacrifice the understanding of the global context.
[0133] In the embodiments of the present disclosure and other possible embodiments, performing a depth enhancement fusion process on the multiple detail extraction features and the convolutional input features to obtain depth detail fusion features includes: performing dilated convolution and attention processing on a part of the multiple detail extraction features to obtain corresponding first fusion sub-features; performing attention processing on another part of the multiple detail extraction features to obtain corresponding second fusion sub-features; performing a concatenation process on the first fusion sub-features and the convolutional input features to obtain a first concatenated feature; and performing an addition process on the first concatenated feature and the second fusion sub-features to obtain the depth detail fusion features.
[0134] In embodiments of the present disclosure and other possible embodiments, depth enhancement fusion processing may be performed using a feature enhancement and fusion module. Among them, dilated convolutions may be respectively performed on the first detail extraction feature, the second detail extraction feature, the third detail extraction feature, and the fourth detail extraction feature, and correspondingly, the first dilated convolution input feature, the second dilated convolution input feature, the third dilated convolution input feature, and the fourth dilated convolution input feature are obtained. Then, attention feature extraction may be respectively performed on the first to fourth dilated convolution input features to obtain four corresponding first fusion sub-features. Specifically, it may be implemented through the attention mechanism CBAM, and the present disclosure does not make specific limitations thereto. In addition, convolutional attention feature extraction may be directly performed on the fifth detail extraction feature to obtain a corresponding second fusion sub-feature. Features of different sizes and directions are extracted through a combination of multi-scale convolutional kernels, and a fusion strategy combining an adaptive attention mechanism is used to obtain corresponding fusion sub-features. After the fusion sub-features obtained on different branches are obtained, further fusion enhancement may be performed. Among them, the four first fusion sub-features and the convolutional input feature may be subjected to a concatenation process to obtain a first concatenated feature; then, after performing a 1×1 convolution on the first concatenated feature, it is added and fused with the second fusion sub-feature to implement the addition process of the first concatenated feature and the second fusion sub-feature, and the depth detail fusion feature is obtained.
[0135] In embodiments of the present disclosure and other possible embodiments, in the case where the depth detail fusion feature is obtained, edge feature-guided fusion may be further performed to improve the detail recognition ability of edge features.
[0136] In embodiments of the present disclosure and other possible embodiments, performing edge guidance processing on the depth detail fusion feature to obtain an edge detail enhancement feature includes: performing an activation process on the depth detail fusion feature to obtain a first depth fusion activation feature; performing multi-scale attention processing on the first depth fusion activation feature to obtain an attention feature; performing a residual convolution operation on the attention feature to obtain the edge detail enhancement feature.
[0137] In the embodiments of the present disclosure and other possible embodiments, the Relu activation function is applied to the depth detail fusion features to obtain the first depth fusion activation features. After obtaining the first depth fusion activation features, a multi-scale attention module can be used to perform multi-scale attention processing to obtain corresponding attention features, and then residual processing is performed. The residual processing in the embodiments of the present disclosure may include performing a convolution operation on the attention features using a 3×3 convolutional kernel, adding the obtained convolution input features to the attention features to obtain the features after residual processing. Then, the features after residual processing can be subjected to a 3×3 convolution operation, adaptive average pooling, convolution, and sigmoid activation processing, and the result is multiplied by the features output by the residual processing to obtain the edge detail enhancement features. The edge detail enhancement features can be used as the first detail enhancement feature, the second detail enhancement feature, or the third detail enhancement feature.
[0138] In the embodiments of the present disclosure and other possible embodiments, in terms of edge-guided feature fusion, the embodiments of the present disclosure can further integrate such edge information into the feature fusion stage and construct a context attention mechanism based on edge reinforcement. This mechanism can dynamically adjust the edge weights in the feature fusion process, thereby highlighting the boundary features. Element-level weighted and multiplication operations are also introduced. By embedding the edge information into the entire feature extraction process, the robustness to complex camouflage scenarios is significantly improved. Specifically, the result output by the feature enhancement and fusion part is superimposed after passing through a 3×3 convolutional layer, and then multiplied by the result passing through a 3×3 convolutional layer and a setting module, and finally the entire result is output. Such a design can explicitly extract edge information, integrate it into feature fusion, strengthen the model's attention to the edge region, and in the feature fusion process, introducing edge information can supplement the information that may be missing in the local features, thereby enhancing the model's ability to capture the subtle differences in the region. Among them, the setting module includes: an adaptive average pooling operation, a convolutional kernel Sigmoid activation function.
[0139] In the embodiments of the present disclosure and other possible embodiments, feature maps with different resolutions can be synergistically fused to obtain a common feature fusion map. The input of the synergistic fusion process is configured as a second detail-enhanced input feature and a third detail-enhanced input feature. The synergistic fusion process includes: respectively performing attention processing on the second detail-enhanced input feature and the third detail-enhanced input feature to correspondingly obtain a first attention input feature and a second attention input feature; performing a multi-scale convolution operation on the first attention input feature to obtain a multi-scale convolution input feature; performing pooling processing on the second attention input feature, and performing the multi-scale convolution operation on the result of the pooling processing to obtain a multi-scale pooled convolution input feature; obtaining a synergistic multi-scale fusion feature corresponding to the second detail-enhanced input feature and the third detail-enhanced input feature based on the multi-scale convolution input feature and the multi-scale pooled convolution input feature.
[0140] In the embodiments of the present disclosure and other possible embodiments, the input of the synergistic fusion process may include two parts of features, such as the first feature and the second feature, the second feature and the third feature, the third feature and the fourth feature, any combination of a low-level multi-scale fusion feature and a high-level fusion feature. In the above combinations, the features can be divided into high-variation-rate features and low-resolution features. In the synergistic fusion process, attention processing can be respectively performed on the second detail-enhanced input feature (high-resolution feature) and the third detail-enhanced input feature (low-resolution feature) to correspondingly obtain a first attention input feature and a second attention input feature; performing a multi-scale convolution operation on the first attention input feature to obtain a multi-scale convolution input feature; performing pooling processing on the second attention input feature, and performing the multi-scale convolution operation on the result of the pooling processing to obtain a multi-scale pooled convolution input feature; obtaining a synergistic multi-scale fusion feature corresponding to the second detail-enhanced input feature and the third detail-enhanced input feature based on the multi-scale convolution input feature and the multi-scale pooled convolution input feature.
[0141] In the embodiments of the present disclosure and other possible embodiments, during the synergistic fusion feature fusion process, weighted processing is respectively performed on the low-resolution low-level multi-scale fusion feature and the high-resolution high-level multi-scale fusion feature to obtain the weighted low-level multi-scale fusion feature and high-level multi-scale fusion feature. This process is achieved through a multi-scale channel attention mechanism to ensure effective weighting of features with different resolutions.
[0142] In the embodiments of the present disclosure and other possible embodiments, local and global attention mechanisms are used to weight the low-level multi-scale fusion features and high-level multi-scale fusion features respectively to generate attention-weighted features. These weighted features are fused through convolution operations, and the interpolation upsampling technique is used to adjust the low-resolution low-level multi-scale fusion features to the same scale as the high-resolution high-level multi-scale fusion features, and finally the fused common features are output. In this process, the attention mechanism not only enhances the feature representation of key regions, but also weights different-resolution features using multi-scale channel attention to ensure effective fusion in both spatial and channel dimensions, further improving the perception ability and robustness.
[0143] In the embodiments of the present disclosure and other possible embodiments, after obtaining multiple collaborative multi-scale fusion features, the collaborative multi-scale fusion features can be further optimized. The collaborative multi-scale fusion features are respectively optimized to obtain multiple optimized collaborative multi-scale fusion features, including: for each collaborative multi-scale fusion feature, first perform a convolution operation on it using a 3×3 convolution kernel, which can efficiently extract local features while further reducing the dimensional redundancy of the features. Subsequently, the convolved features are scale-unified through normalization processing to eliminate the numerical differences between different features, enhance the numerical stability of the model, and accelerate convergence. Finally, the normalized features are non-linearly transformed through the ReLU activation function, introducing non-linear factors to enhance the expressive ability of the features in detail, while avoiding the problem of gradient disappearance, and further improving the discriminability and adaptability of the features. Through this series of operations, the collaborative multi-scale fusion features are respectively optimized, and finally multiple optimized features are obtained.
[0144] In the embodiments of the present disclosure and other possible embodiments, after obtaining the optimized features, the multiple optimized collaborative multi-scale fusion features can be subjected to summation processing, and the detection is performed using the result of the summation processing. In the embodiments of the present disclosure, a mask map representing the position of the driving vehicle on the preset navigation route can be generated. For example, sigmoid activation processing can be performed on the result of the summation processing to obtain the detection result of the driving vehicle on the preset navigation route.
[0145] In the embodiments of the present disclosure and other possible embodiments, the obtained result of the summation process can also be used to determine the corresponding image region, extract the radiomics features of the region, and perform a convolution operation on the radiomics features to obtain a omics feature map with the same image scale at multiple moments corresponding to the video stream of the driving vehicle on the preset navigation route. Then, a convolution operation is performed on the omics feature map and the result of the summation process to obtain an optimized detection result. In the embodiments of the present disclosure, a mask map representing the position of the driving vehicle on the preset navigation route can be generated. For example, a sigmoid activation process can be performed on the result of the convolution operation to obtain the detection result of the driving vehicle on the preset navigation route.
[0146] Step S103: Based on the real-variable topological network, predict the driving road conditions of the first set driving vehicle on the preset navigation route.
[0147] In the embodiments of the present disclosure, the method for predicting the driving road conditions corresponding to the preset navigation route based on the real-variable topological network includes: based on the real-variable topological network, determine multiple core nodes corresponding to the intersections in the driving direction on the preset navigation route; respectively predict whether multiple second set vehicles in the real-variable topological network will interfere with the first set driving vehicle at the multiple core nodes; if interference occurs, then based on the cumulative interference time of the multiple second set vehicles and the set time, predict the congestion condition of the corresponding intersection in the real-variable topological network.
[0148] In the embodiments of the present disclosure, the method for respectively predicting whether multiple second set vehicles in the real-variable topological network will interfere with the first set driving vehicle at the multiple core nodes includes: respectively calculate multiple first distances and multiple second distances between the first set driving vehicle and multiple second set vehicles in the real-variable topological network and the multiple core nodes; predict multiple first speeds at which the first set driving vehicle will reach the multiple core nodes corresponding to the remaining driving distance; respectively predict multiple second speeds at which multiple second set vehicles will reach the multiple core nodes corresponding to the multiple second set vehicles; based on the multiple first distances and the multiple first speeds corresponding to the first set driving vehicle, the multiple second distances corresponding to the multiple second set vehicles and the multiple second speeds, determine whether each second set vehicle in the multiple second set vehicles will interfere with the first set driving vehicle passing through the corresponding core node.
[0149] In an embodiment of the present disclosure, the method for determining whether each of the plurality of second preset vehicles obstructs the first preset driving vehicle from passing through the corresponding core node based on the plurality of first distances corresponding to the first preset driving vehicle, the plurality of first speeds, the plurality of second distances corresponding to the plurality of second preset vehicle distances, and the plurality of second speeds includes: determining a plurality of first positions corresponding to the first preset driving vehicle based on the plurality of first distances corresponding to the first preset driving vehicle and the plurality of first speeds; determining a plurality of second positions corresponding to each of the plurality of second preset vehicles based on the plurality of second distances corresponding to the plurality of second preset vehicle distances and the plurality of second speeds; and determining whether the plurality of second preset vehicles obstruct the first preset driving vehicle at the plurality of core nodes at the intersection corresponding to each core node based on the plurality of first positions corresponding to the first preset driving vehicle and the plurality of second positions corresponding to each of the plurality of second preset vehicles.
[0150] In an embodiment of the present disclosure and other possible embodiments, the method for determining whether the plurality of second preset vehicles obstruct the first preset driving vehicle at the plurality of core nodes at the intersection corresponding to each core node based on the plurality of first positions corresponding to the first preset driving vehicle and the plurality of second positions corresponding to each of the plurality of second preset vehicles includes: determining whether each of the plurality of second preset vehicles is respectively in front of the first preset driving vehicle on the driving road corresponding to the intersection corresponding to each core node according to the plurality of first positions corresponding to the first preset driving vehicle and the plurality of second positions corresponding to each of the plurality of second preset vehicles; if so, the plurality of second preset vehicles obstruct the first preset driving vehicle at the plurality of core nodes; otherwise, the plurality of second preset vehicles do not obstruct the first preset driving vehicle at the plurality of core nodes.
[0151] In an embodiment of the present disclosure and other possible embodiments, the method for predicting the congestion situation of the corresponding intersection in the time-varying topology network based on the cumulative obstruction time of the plurality of second preset vehicles includes: predicting the obstruction time corresponding to each second preset vehicle, and determining the cumulative obstruction time according to the obstruction time; if the cumulative obstruction time is less than a first preset time, predicting that the corresponding intersection in the time-varying topology network is not congested; otherwise, if the cumulative obstruction time is less than a second preset time, predicting that the corresponding intersection in the time-varying topology network is slightly congested; otherwise, predicting that the corresponding intersection in the time-varying topology network is severely congested.
[0152] In the embodiments of the present disclosure and other possible embodiments, if the cumulative obstruction time is less than the first set time, it is predicted that the corresponding intersection in the real-time topology network is not congested; if the cumulative obstruction time is greater than the first set time and less than the second set time, it is predicted that the corresponding intersection in the real-time topology network is slightly congested; if the cumulative obstruction time is greater than the second set time, it is predicted that the corresponding intersection in the real-time topology network is severely congested.
[0153] In the embodiments of the present disclosure, the method for predicting the multiple first speeds at which the first set driving vehicle reaches the multiple core nodes in the remaining driving distance includes: obtaining the distance already traveled, the time already traveled, and the road conditions corresponding to the remaining driving distance of the first set driving vehicle on the preset navigation route; wherein, the road conditions include one or several of the number of the first lanes corresponding to the driving direction of the first set driving vehicle, the number of the first vehicles corresponding to each lane, and the weather condition; based on the distance already traveled, the time already traveled, and the road conditions corresponding to the remaining driving distance of the first set driving vehicle on the preset navigation route, using a machine learning prediction model or a time series analysis algorithm, respectively predicting the multiple first speeds at which the first set driving vehicle reaches the multiple core nodes in the remaining driving distance.
[0154] In the embodiments of the present disclosure, the method for respectively predicting the multiple second speeds at which multiple second set vehicles reach the multiple core nodes includes: obtaining the distance already traveled, the time already traveled, and the road conditions corresponding to the remaining driving distance of the multiple second set driving vehicles on the preset navigation route; wherein, the road conditions include one or several of the number of the second lanes corresponding to the driving directions of the multiple second set driving vehicles towards the multiple core nodes, the number of the second vehicles corresponding to each lane, and the weather condition; based on the distance already traveled, the time already traveled, and the road conditions corresponding to the remaining driving distance of the second set driving vehicles on the preset navigation route, using a machine learning prediction model or a time series analysis algorithm, respectively predicting the multiple second speeds at which the first set driving vehicle reaches the multiple core nodes in the remaining driving distance.
[0155] In the embodiments of the present disclosure and other possible embodiments, the machine learning prediction model can be configured as one or several corresponding prediction models among a support vector machine, a decision tree, a random forest, a K-nearest neighbor, a logistic regression, an adaptive boosting, a linear discriminant analysis, and a multi-layer perceptron.
[0156] In an embodiment of the present disclosure, the road condition prediction method further includes: if it is predicted that the driving road condition corresponding to the first set driving vehicle on the preset navigation route is congested, adjusting the current starting point of the first set driving vehicle to an adjusted navigation route corresponding to the preset end point; wherein, the congestion is configured as slight congestion or severe congestion; respectively determining a plurality of intersection roads connected to the adjusted preset navigation route; in real time, with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the adjusted navigation route as the radius, constructing a real-time variable topology network corresponding to the intersection roads within the radius among the plurality of intersection roads; and predicting the driving road condition corresponding to the first set driving vehicle on the adjusted navigation route based on the real-time variable topology network.
[0157] In an embodiment of the present disclosure and other possible embodiments, the constructing, in real time, with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the adjusted navigation route as the radius, a real-time variable topology network corresponding to the intersection roads within the radius among the plurality of intersection roads includes: in real time, with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the adjusted navigation route as the radius, determining a real-time variable topology range; within the real-time variable topology range, using the first set driving vehicle and the intersection roads within the radius as nodes, and using the route corresponding to the intersection road closest to the first set driving vehicle in the adjusted navigation route and the routes between the intersection roads within the radius as edges, constructing a real-time variable topology network corresponding to the intersection roads within the radius among the plurality of intersection roads.
[0158] In an embodiment of the present disclosure and other possible embodiments, it further includes: optimizing the real-time variable topology network, and the optimization method includes: deleting the passed nodes of the first set driving vehicle in the real-time variable topology network and the edges corresponding to the nodes connected to the passed nodes and not on the adjusted navigation route and their routes.
[0159] In an embodiment of the present disclosure and other possible embodiments, during the process of constructing, within the real-time variable topology range, a real-time variable topology network corresponding to the intersection roads within the radius among the plurality of intersection roads with the first set driving vehicle and the intersection roads within the radius as nodes and the route corresponding to the intersection road closest to the first set driving vehicle in the adjusted navigation route and the routes between the intersection roads within the radius as edges, if there is a route in the direction of the first set driving vehicle among the routes between the intersection roads within the radius, the edge corresponding to the real-time variable topology network is retained; otherwise, the edge corresponding to the real-time variable topology network is deleted.
[0160] In the embodiments of the present disclosure and other possible embodiments, when deleting the edges corresponding to the real-time variable topological network, there is no route in the direction of the first set driving vehicle between the intersections within the radius in the direction of the first set driving vehicle (there is no road from the intersection to the first set driving vehicle, i.e., a one-way street).
[0161] In the embodiments of the present disclosure and other possible embodiments, in summary, within the real-time variable topological range, taking the first set driving vehicle and the intersections within the radius as nodes, and taking the route corresponding to the nearest intersection of the first set driving vehicle in the adjusted navigation route and the route in the direction of the first set driving vehicle between the intersections within the radius as edges, a real-time variable topological network corresponding to the intersections within the radius of the multiple intersections is constructed.
[0162] In the embodiments of the present disclosure and other possible embodiments, before constructing the real-time variable topological network corresponding to the intersections within the radius of the multiple intersections, the method for determining the first set driving vehicle corresponding to the adjusted navigation route in real time includes: obtaining multiple moment images corresponding to the video stream of the first set driving vehicle on the adjusted navigation route in real time; using an object detection model to detect the license plates of all driving vehicles in the multiple moment images to determine the first set driving vehicle corresponding to the adjusted navigation route.
[0163] In the embodiments of the present disclosure and other possible embodiments, the method for using an object detection model to detect the license plates of all set driving vehicles in the multiple moment images to determine the first set driving vehicle corresponding to the adjusted navigation route includes: using the object detection model to detect the license plates of all driving vehicles in the multiple moment images to obtain corresponding multiple license plate detection frames; respectively identifying the license plate numbers corresponding to the multiple license plate detection frames to obtain corresponding multiple identified license plate numbers; based on the fact that the multiple identified license plate numbers are consistent with the preset license plate number corresponding to the first set driving vehicle of the adjusted navigation route, the driving vehicle corresponding to the consistent preset license plate number is determined as the first set driving vehicle corresponding to the adjusted navigation route.
[0164] In embodiments of the present disclosure and other possible embodiments, if the license plates of all driving vehicles in the multi-moment images are detected by using the target detection model, and the first set driving vehicle corresponding to the preset license plate number for adjusting the navigation route is not determined, the multi-scale features of the multi-moment images are respectively obtained to obtain corresponding high-level multi-scale fusion features and low-level multi-scale fusion features; the grouping detail enhancement and fusion processing are respectively performed on the multi-scale features, high-level multi-scale fusion features and low-level multi-scale fusion features to obtain corresponding multiple collaborative multi-scale fusion features; the collaborative multi-scale fusion features are respectively optimized to obtain corresponding multiple optimized collaborative multi-scale fusion features; the first set driving vehicle in the multi-moment images is respectively detected or determined based on the multiple optimized collaborative multi-scale fusion features.
[0165] In embodiments of the present disclosure and other possible embodiments, the method for predicting the driving road conditions corresponding to the adjusted navigation route based on the real-variable topological network includes: determining, based on the real-variable topological network, a plurality of core nodes corresponding to the intersection in the driving direction in the adjusted navigation route; respectively predicting whether a plurality of second set vehicles in the real-variable topological network obstruct the first set driving vehicle at the plurality of core nodes; if an obstruction occurs, predicting the congestion condition of the corresponding intersection in the real-variable topological network based on the cumulative obstruction time of the plurality of second set vehicles and the set time.
[0166] In embodiments of the present disclosure and other possible embodiments, the method for respectively predicting whether a plurality of second set vehicles in the real-variable topological network obstruct the first set driving vehicle at the plurality of core nodes includes: respectively calculating a plurality of first distances and a plurality of second distances between the first set driving vehicle and the plurality of second set vehicles in the real-variable topological network from the plurality of core nodes; predicting a plurality of first speeds at which the first set driving vehicle reaches the plurality of core nodes corresponding to the remaining driving distance; respectively predicting a plurality of second speeds at which the plurality of second set vehicles reach the plurality of core nodes corresponding to the plurality of second set vehicles; determining whether each second set vehicle in the plurality of second set vehicles obstructs the first set driving vehicle from passing through the corresponding core node based on the plurality of first distances and the plurality of first speeds corresponding to the first set driving vehicle, the plurality of second distances corresponding to the plurality of second set vehicles, and the plurality of second speeds.
[0167] In the embodiments of the present disclosure and other possible embodiments, the method for determining whether each of the plurality of second preset vehicles obstructs the first preset driving vehicle from passing through the corresponding core node based on the plurality of first distances corresponding to the first preset driving vehicle, the plurality of first speeds, the plurality of second distances corresponding to the plurality of second preset vehicle distances, and the plurality of second speeds includes: determining a plurality of first positions corresponding to the first preset driving vehicle based on the plurality of first distances corresponding to the first preset driving vehicle and the plurality of first speeds; determining a plurality of second positions corresponding to each of the plurality of second preset vehicles based on the plurality of second distances corresponding to the plurality of second preset vehicle distances and the plurality of second speeds; and determining whether the plurality of second preset vehicles obstruct the first preset driving vehicle at the plurality of core nodes at the intersection corresponding to each core node based on the plurality of first positions corresponding to the first preset driving vehicle and the plurality of second positions corresponding to each of the plurality of second preset vehicles.
[0168] In the embodiments of the present disclosure and other possible embodiments, the method for determining whether the plurality of second preset vehicles obstruct the first preset driving vehicle at the plurality of core nodes at the intersection corresponding to each core node based on the plurality of first positions corresponding to the first preset driving vehicle and the plurality of second positions corresponding to each of the plurality of second preset vehicles includes: determining whether each of the plurality of second preset vehicles is respectively in front of the first preset driving vehicle on the driving road corresponding to the intersection corresponding to each core node according to the plurality of first positions corresponding to the first preset driving vehicle and the plurality of second positions corresponding to each of the plurality of second preset vehicles; if so, the plurality of second preset vehicles obstruct the first preset driving vehicle at the plurality of core nodes; otherwise, the plurality of second preset vehicles do not obstruct the first preset driving vehicle at the plurality of core nodes.
[0169] In the embodiments of the present disclosure and other possible embodiments, the method for predicting the congestion situation of the corresponding intersection in the time-varying topological network based on the cumulative obstruction time of the plurality of second preset vehicles includes: predicting the obstruction time corresponding to each second preset vehicle, and determining the cumulative obstruction time according to the obstruction time; if the cumulative obstruction time is less than the first preset time, predicting that the corresponding intersection in the time-varying topological network is not congested; otherwise, if the cumulative obstruction time is less than the second preset time, predicting that the corresponding intersection in the time-varying topological network is slightly congested; otherwise, predicting that the corresponding intersection in the time-varying topological network is severely congested.
[0170] In the embodiments of the present disclosure and other possible embodiments, if the cumulative obstruction time is less than the first set time, it is predicted that the corresponding intersection in the real-time variable topology network is not congested; if the cumulative obstruction time is greater than the first set time and less than the second set time, it is predicted that the corresponding intersection in the real-time variable topology network is slightly congested; if the cumulative obstruction time is greater than the second set time, it is predicted that the corresponding intersection in the real-time variable topology network is severely congested.
[0171] In the embodiments of the present disclosure and other possible embodiments, the method for predicting the multiple first speeds corresponding to the multiple core nodes when the first set driving vehicle reaches the remaining driving distance includes: obtaining the traveled distance, traveled time, and road conditions corresponding to the remaining driving distance of the first set driving vehicle on the adjusted navigation route; wherein, the road conditions include one or several of the number of the first lanes corresponding to the driving direction of the first set driving vehicle, the number of the first vehicles corresponding to each lane, and the weather condition; based on the traveled distance, traveled time, and road conditions corresponding to the remaining driving distance of the first set driving vehicle on the adjusted navigation route, using a machine learning prediction model or a time series analysis algorithm to respectively predict the multiple first speeds corresponding to the multiple core nodes when the first set driving vehicle reaches the remaining driving distance.
[0172] In the embodiments of the present disclosure and other possible embodiments, the method for respectively predicting the multiple second speeds corresponding to the multiple core nodes when multiple second set vehicles reach includes: obtaining the traveled distance, traveled time, and road conditions corresponding to the remaining driving distance of the multiple second set driving vehicles on the adjusted navigation route; wherein, the road conditions include one or several of the number of the second lanes corresponding to the driving directions of the multiple second set driving vehicles towards the multiple core nodes, the number of the second vehicles corresponding to each lane, and the weather condition; based on the traveled distance, traveled time, and road conditions corresponding to the remaining driving distance of the second set driving vehicles on the adjusted navigation route, using a machine learning prediction model or a time series analysis algorithm to respectively predict the multiple second speeds corresponding to the multiple core nodes when the first set driving vehicle reaches the remaining driving distance.
[0173] The execution subject of the road condition prediction method based on computer vision technology can be a road condition prediction device or system based on computer vision technology. For example, the road condition prediction method based on computer vision technology can be executed by a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the road condition prediction method based on computer vision technology can be implemented by a processor calling computer-readable instructions stored in a memory.
[0174] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0175] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: an acquisition unit, configured to acquire a preset navigation route corresponding to a preset starting point and a preset ending point of a first set driving vehicle, and respectively determine a plurality of intersection roads connected to the preset navigation route; a construction unit, configured to construct, in real time, a real-time variable topological network corresponding to the intersection roads within the radius with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius; a prediction unit, configured to predict the driving road condition of the first set driving vehicle on the preset navigation route based on the real-time variable topological network.
[0176] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: an electronic device, the electronic device being configured with a processor and a memory for storing processor-executable instructions; wherein, the processor is configured to call the instructions stored in the memory to perform the above road condition prediction method.
[0177] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: a computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions, when executed by a processor, implement the above road condition prediction method.
[0178] According to one aspect of the present disclosure, there is provided a road condition prediction device or system based on computer vision technology, including: a computer program product, on which a computer program / instructions are provided, and the computer program / instructions, when executed by a processor, implement the above road condition prediction method.
[0179] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A road condition prediction method based on computer vision technology, characterized in that: include: Obtaining a preset navigation route corresponding to a preset starting point and a preset end point of a first set driving vehicle, and respectively determining a plurality of intersections connected to the preset navigation route; In real time, taking the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, constructing a real-time topological network corresponding to the intersections within the radius at the plurality of intersections; Based on the real-time topology network, a driving condition corresponding to the preset navigation route of the first set driving vehicle is predicted.
2. The road condition prediction method according to claim 1, characterized in that: The real-time topological network corresponding to the intersections within the radius of the plurality of intersections is constructed with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, including: In real time, taking the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius, determining the real change topology range; Within the scope of the actual topology, the first set driving vehicle and the intersections within the radius are taken as nodes, and the route corresponding to the nearest intersection of the first set driving vehicle in the preset navigation route and the route between the intersections within the radius are taken as edges, to construct a actual topology network corresponding to the intersections within the radius of the multiple intersections.
3. The road condition prediction method according to claim 2, characterized in that: Before constructing the real-time topological network corresponding to the intersections of the plurality of intersections within the radius, the method for determining the first set driving vehicle corresponding to the preset navigation route in real time includes: Acquire in real time multiple-time images corresponding to the video stream of the first set driving vehicle on the preset navigation route; Detecting the license plates of all driving vehicles in the multi-time images using a target detection model to determine the first set driving vehicle corresponding to the preset navigation route; and / or, The method of detecting the license plates of all set driving vehicles in the multi-time images by using the target detection model to determine the first set driving vehicle corresponding to the preset navigation route includes: The target detection model is used to detect the license plates of all driving vehicles in the multi-time images to obtain corresponding multiple license plate detection frames; Respectively identifying the corresponding license plate numbers in the plurality of license plate detection frames to obtain a plurality of corresponding identified license plate numbers; Based on the fact that the multiple identified license plate numbers are consistent with the preset license plate number corresponding to the first set driving vehicle of the preset navigation route, the driving vehicle corresponding to the preset license plate number is determined as the first set driving vehicle corresponding to the preset navigation route.
4. The road condition prediction method according to claim 3, characterized in that: If the target detection model is used to detect the license plates of all driving vehicles in the multi-time images, and the first set driving vehicle corresponding to the preset license plate number of the preset navigation route is not determined, then the corresponding high-level multi-scale fusion features and low-level multi-scale fusion features are obtained for the multiple multi-scale features of the multi-time images respectively; Performing group detail enhancement and fusion processing on the multiple multi-scale features, high-level multi-scale fusion features and low-level multi-scale fusion features respectively to obtain corresponding multiple collaborative multi-scale fusion features; Optimizing the collaborative multi-scale fusion features respectively to obtain a corresponding plurality of optimized collaborative multi-scale fusion features; The first set driving vehicle in the multi-time images is detected or determined based on the multiple optimized collaborative multi-scale fusion features respectively.
5. The road condition prediction method according to any one of claims 1 to 4, characterized in that: The method for predicting the driving road conditions corresponding to the preset navigation route based on the real variable topology network includes: Based on the real-time topological network, determining a plurality of core nodes corresponding to intersections in the travel direction in the preset navigation route; respectively predicting whether a plurality of second-set vehicles in the real-change topology network will cause obstruction to the first-set driving vehicle at the plurality of core nodes; If an obstruction occurs, the congestion condition of the corresponding intersection road in the real-time topology network is predicted based on the accumulated obstruction time and the set time of the plurality of second set vehicles.
6. The road condition prediction method according to claim 5, characterized in that: The method of respectively predicting whether a plurality of second-set vehicles in the real-variable topology network will cause obstruction to the first-set driving vehicle at the plurality of core nodes comprises: Respectively calculating a plurality of first distances and a plurality of second distances between the first set driving vehicle and a plurality of second set vehicles in the actual topology network and the plurality of core nodes; Predicting that the first setting driving vehicle reaches a plurality of first speeds corresponding to the plurality of core nodes in the remaining driving distance; respectively predicting a plurality of second speeds corresponding to a plurality of second set vehicles arriving at the plurality of core nodes; Based on the multiple first distances and the multiple first speeds corresponding to the first set driving vehicle and the multiple second distances and the multiple second speeds corresponding to the multiple second set vehicle distances, determine whether each of the multiple second set vehicles prevents the first set driving vehicle from passing the corresponding core node.
7. The road condition prediction method according to claim 6, characterized in that: The method for determining whether each of the plurality of second setting vehicles prevents the first setting driving vehicle from passing through the corresponding core node based on the plurality of first distances and the plurality of first speeds corresponding to the first setting driving vehicle and the plurality of second distances and the plurality of second speeds corresponding to the plurality of second setting vehicle distances includes: determining a plurality of first positions corresponding to the first setting driving vehicle based on the plurality of first distances and the plurality of first speeds corresponding to the first setting driving vehicle; Determining a plurality of second positions corresponding to each of the plurality of second setting vehicles based on the plurality of second distances used for the plurality of second setting vehicle distance pairs and the plurality of second speeds; Based on the multiple first positions corresponding to the first set driving vehicle and the multiple second positions corresponding to each of the multiple second set vehicles, determine whether the multiple second set vehicles cause obstruction to the first set driving vehicle at the multiple core nodes at the intersection corresponding to each core node.
8. The road condition prediction method according to any one of claims 6 or 7, characterized in that: The method for predicting that the first set driving vehicle reaches a plurality of first speeds corresponding to the plurality of core nodes in the remaining driving distance includes: Acquire the road conditions corresponding to the distance traveled, the travel time, and the remaining travel distance of the first set driving vehicle on the preset navigation route; Based on the distance traveled by the first set driving vehicle on the preset navigation route, the travel time, and the road conditions corresponding to the remaining travel distance, a machine learning prediction model or a time series analysis algorithm is used to respectively predict a plurality of first speeds corresponding to the plurality of core nodes reached by the first set driving vehicle in the remaining travel distance; and / or The method for respectively predicting that a plurality of second set vehicles arrive at a plurality of second speeds corresponding to the plurality of core nodes comprises: Acquire a plurality of road conditions corresponding to the travel distance, travel time, and remaining travel distance of the second set driving vehicle on the preset navigation route; Based on the road conditions corresponding to the distance traveled, the driving time, and the remaining driving distance of the second-set driving vehicle on the preset navigation route, a machine learning prediction model or a time series analysis algorithm is used to predict multiple second speeds corresponding to the first-set driving vehicle reaching the multiple core nodes in the remaining driving distance.
9. The road condition prediction method according to any one of claims 1 to 8, characterized in that: Also includes: If it is predicted that the driving road condition corresponding to the preset navigation route of the first set driving vehicle is congested, adjusting the navigation route corresponding to the current starting point of the first set driving vehicle to the preset end point; wherein the congestion is configured as slight congestion or severe congestion; Respectively determine and adjust multiple intersections connected to a preset navigation route; In real time, taking the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the adjusted navigation route as the radius, constructing a real topological network corresponding to the intersections within the radius at the plurality of intersections; Based on the real-time topology network, a driving condition corresponding to the adjusted navigation route of the first set driving vehicle is predicted.
10. A road condition prediction system based on computer vision technology, characterized in that: include: An acquisition unit, used to acquire a preset navigation route corresponding to a preset starting point and a preset end point of a first set driving vehicle, and respectively determine a plurality of intersections connected to the preset navigation route; A construction unit is used to construct a real-time topological network corresponding to the intersections within the radius at the plurality of intersections, with the first set driving vehicle as the center and the real-time remaining driving distance corresponding to the preset navigation route as the radius; A prediction unit, configured to predict, based on the real-time topology network, a driving condition corresponding to the preset navigation route of the first set driving vehicle; or, The invention comprises: an electronic device, wherein the electronic device is configured with a processor and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the road condition prediction method according to any one of claims 1 to 9; or, The method comprises: a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the road condition prediction method according to any one of claims 1 to 9; or The invention comprises: a computer program product, wherein a computer program / instruction is set on the computer program product, and when the computer program / instruction is executed by a processor, the road condition prediction method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Refined path planning method based on road network rasterized road traffic flow prediction
CN112629533A
Visualizing unidirectional traffic information
US20180252548A1
Smart navigation method and system based on topological map
US20210302585A1
Map Data Processing Method and Apparatus
US20240410708A1
Map matching method and device based on trajectory topology
WO2024192788A1