A traffic flow detection and prediction method based on unmanned aerial vehicle feedback cooperative optimization
By using multi-level modeling and multimodal coding networks for traffic data processing through UAV collaborative optimization, the adaptability and accuracy issues of traffic flow detection and prediction in existing technologies have been resolved, achieving efficient and accurate detection and prediction of traffic flow.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALADDIN UAV (SHENZHEN) CO LTD
- Filing Date
- 2025-05-09
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, methods for traffic flow detection and prediction by collecting image data using fixed cameras are poorly adaptable to complex traffic scenarios and have insufficient prediction accuracy.
A method based on UAV feedback and collaborative optimization is adopted, which uses a multi-level modeling network and a multimodal coding network to analyze and process image and text traffic data, and extract structured traffic flow features and prediction features.
It enables real-time and complete perception of traffic flow, improves the accuracy and effectiveness of detection and prediction, and provides reliable data support for traffic scheduling.
Smart Images

Figure CN120449100B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of traffic flow detection technology, and in particular to a traffic flow detection and prediction method based on UAV feedback collaborative optimization. Background Technology
[0002] In the field of traffic flow optimization technology, there is a need to detect and predict traffic flow in order to recommend traffic driving strategies to drivers.
[0003] In related technologies, fixed cameras are deployed to collect image data, and traditional image processing algorithms are used to estimate indicators such as traffic density and speed, thereby realizing the identification, detection and trend prediction of the current traffic operation status. However, such methods have the disadvantages of limited ability to express data features and insufficient characterization of temporal evolution, resulting in poor adaptability of detection and prediction results to complex traffic scenarios and insufficient prediction accuracy. Summary of the Invention
[0004] Therefore, it is necessary to address the aforementioned technical problems by providing a method, apparatus, computer equipment, and computer-readable storage medium for traffic flow detection and prediction based on UAV feedback collaborative optimization, in order to improve the accuracy and effectiveness of traffic flow detection and prediction, and provide reliable data support for traffic scheduling and control.
[0005] Firstly, this application provides a traffic flow detection and prediction method based on UAV feedback collaborative optimization, including:
[0006] In the current traffic environment, current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in historical time periods are acquired. The image capturing device consists of multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment.
[0007] Based on a preset multi-level modeling network, the current image traffic data is parsed and processed to obtain the current traffic flow characteristics. Based on the current traffic flow characteristics, the traffic flow detection result of the current traffic environment in the current time period is generated.
[0008] Based on a preset multimodal coding network, the current image traffic data and the historical text traffic data are parsed and processed to obtain predicted traffic flow features. Based on the predicted traffic flow features, a traffic flow prediction result for the current traffic environment in the future time period is generated.
[0009] Secondly, this application also provides a traffic flow detection and prediction device based on UAV feedback collaborative optimization, comprising:
[0010] The acquisition module is used to acquire current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in the historical time period in the current traffic environment. The image capturing device is a combination of multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment.
[0011] The detection module is used to perform data parsing processing on the current image traffic data based on a preset multi-level modeling network to obtain the current traffic flow characteristics, and generate the traffic flow detection result of the current traffic environment in the current time period based on the current traffic flow characteristics.
[0012] The prediction module is used to perform data parsing and processing on the current image traffic data and the historical text traffic data based on a preset multimodal coding network to obtain predicted traffic flow features, and generate a traffic flow prediction result for the current traffic environment in the future time period based on the predicted traffic flow features.
[0013] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the above steps.
[0014] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the above steps.
[0015] The aforementioned traffic flow detection and prediction method, apparatus, computer equipment, and computer-readable storage medium based on UAV feedback collaborative optimization firstly, comprehensively perceives the current traffic environment by using current image traffic data collected by an image-capturing device in the current time period and historical text traffic data stored in historical time periods, thereby ensuring the real-time nature and data integrity of subsequent traffic flow detection and prediction. Secondly, multi-level modeling processing is performed on the current image traffic data using a multi-level modeling network, thereby extracting structured current traffic flow features and generating traffic flow detection results that can be used for quantitative analysis. Thirdly, multi-modal coding processing is performed on the current image traffic data and historical text traffic data using a multi-modal coding network, thereby extracting predictive traffic flow features that reflect the changing trends of traffic operation status and generating traffic flow prediction results that can be used for quantitative analysis. Based on this, through a continuous processing chain of data acquisition, traffic flow detection, and traffic flow prediction, quantitative perception and trend reasoning capabilities of the current traffic environment are achieved, improving the accuracy and effectiveness of traffic flow detection and prediction, and providing reliable data support for traffic scheduling and control. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a traffic flow detection and prediction method based on UAV feedback collaborative optimization in one embodiment.
[0018] Figure 2 This is a schematic diagram of a process for generating current traffic flow characteristics based on a multi-level modeling network in one embodiment;
[0019] Figure 3 This is a schematic diagram of the process for generating current traffic flow characteristics based on a multi-level modeling network in another embodiment;
[0020] Figure 4 This is a schematic diagram of the process of generating neck structure output based on semantic relationship modeling layer in a multi-level modeling network in one embodiment;
[0021] Figure 5 This is a schematic diagram of a road traffic flow detection algorithm based on real-time UAV data and multi-level relationship modeling in one embodiment.
[0022] Figure 6 This is a schematic diagram of a process for generating predicted traffic flow features based on a multimodal coding network in one embodiment;
[0023] Figure 7 This is a schematic diagram illustrating the process of generating historical traffic flow text features by combining a text encoding layer and a bidirectional long and short time series modeling layer in a multimodal coding network in one embodiment.
[0024] Figure 8 This is a schematic diagram of the structure of a future road traffic flow prediction algorithm based on joint optimization of historical text traffic data and real-time feedback from drones in one embodiment.
[0025] Figure 9 This is a schematic diagram illustrating the process of generating aligned data corresponding to each data layer based on the text visual coding layer in a multimodal coding network in one embodiment.
[0026] Figure 10 This is a schematic diagram of the data alignment algorithm based on the text visual data alignment module in one embodiment;
[0027] Figure 11 This is a structural block diagram of a traffic flow detection and prediction device based on UAV feedback collaborative optimization in one embodiment. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] In one embodiment, such as Figure 1 As shown, a traffic flow detection and prediction method based on UAV feedback collaborative optimization is provided. This embodiment illustrates the method by applying it to a server. It is understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S101 to S103.
[0030] Step S101: In the current traffic environment, acquire current image traffic data collected by preset image capturing devices in the current time period and historical text traffic data stored in historical time periods. The image capturing devices are multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment.
[0031] The current traffic environment refers to the traffic operation status within a specified area, which describes the comprehensive traffic conditions such as vehicles, pedestrians, road facilities, and traffic light control duration within that specified area.
[0032] Among them, image capturing equipment refers to equipment deployed in the current traffic environment to obtain real-time traffic images, that is, traffic data in the form of image data related to factors such as vehicles, pedestrians, and road facilities; image capturing equipment can refer to tethered drone satellites hovering in the air, which are equipped with high-pixel cameras and can achieve centimeter-level real-time resolution of ground objects.
[0033] Among them, current image traffic data refers to image traffic data collected by image capturing equipment during the current time period, reflecting the traffic operation status during the current time period; historical text traffic data refers to text traffic data stored in text format, reflecting the traffic operation status during historical time periods.
[0034] For example, in a city traffic map, the physical space corresponding to a certain map area can be designated as the current traffic environment. For instance, the physical space corresponding to the entire city in the city traffic map can be used as the current traffic environment, or the physical spaces corresponding to multiple interconnected main roads in the city traffic map can be used as the current traffic environment. Furthermore, over time, the network model or algorithm used for traffic flow detection and prediction in this embodiment will generate a large amount of historical data. This historical data reflects the traffic conditions of various roads in the current traffic environment at different historical moments, and this historical data is stored in a database in text form, serving as historical text traffic data for the historical time period in the scenario of detecting and predicting traffic flow in the current time period.
[0035] The image capturing equipment refers to tethered drones deployed on a city-by-city basis. This means that several tethered drones are evenly distributed over the city, and their camera units continuously acquire centimeter-level precision image data in real time, achieving comprehensive visual coverage and global traffic monitoring of the entire city's traffic area. The number, distribution, and spatial distances of the tethered drones deployed over the city are determined by combining factors such as the city's road distribution, boundary outline, area size, and the field of view of the tethered drones. This allows the combined global field of view of a number of tethered drones to comprehensively cover the entire city's traffic area. The field of view of a single tethered drone can be defined as a circular area centered on its location and with the furthest visible distance as its radius.
[0036] Furthermore, the structural composition and corresponding functions of the aforementioned tethered drone can be referenced from a continuously operational internet-connected drone (publication number CN105652886A) in related technologies. This is a tethered drone satellite equipped with a high-pixel camera and a field of view of up to 5 square kilometers. On the one hand, the drone can achieve continuous operation based on continuous power supply, thereby maintaining its working state continuously and without interruption. On the other hand, based on its higher wind and rain resistance, the drone can work continuously under general weather conditions, thereby achieving real-time and uninterrupted acquisition of image data with centimeter-level accuracy.
[0037] Step S102: Based on the preset multi-level modeling network, perform data parsing processing on the current image traffic data to obtain the current traffic flow characteristics, and generate the traffic flow detection results of the current traffic environment in the current time period based on the current traffic flow characteristics.
[0038] Among them, the multi-level modeling network represents a neural network structure used to extract and analyze traffic flow features at different levels from the original traffic data layer by layer. It is used to transform basic visual elements (such as edges, colors, and motion) in the original image data into structured representation features that can reflect the traffic operation status. For example, shallow networks are used to extract vehicle outlines and boundary information, while deep networks are used to infer the overall traffic flow density or traffic congestion trends in local areas.
[0039] Among them, the current traffic flow characteristics represent representative statistical or distributional indicators extracted from the current image traffic data, which are used to describe the traffic operation status in the current time period. For example, the distance between each vehicle and its corresponding destination in the current time period, the number of vehicles passing through each road per unit time, the average vehicle speed, vehicle density, lane saturation, and intersection queue length, etc., and the number and density of pedestrians crossing each road in the current time period, as well as their flow direction and trend, etc.
[0040] Among them, traffic flow detection results represent structured output information that reflects the traffic operation status in the current time period after modeling and processing based on the current image traffic data. It is used to describe key indicators such as the number, distribution, speed or congestion level of traffic flow in the current time period. For example, traffic flow detection results may include the number of vehicles passing through a certain intersection per unit time in the current time period, the average speed of each lane, and whether there are abnormal traffic events in the road segment.
[0041] For example, current image traffic data is input into a multi-level modeling network. Through layer-by-layer feature parsing operations, the low-level visual features in the original image data are gradually transformed into structured information that can characterize traffic operation status. For instance, during the parsing process, elements such as target contours, motion trajectories, and spatial distribution in the original image data are first extracted. Then, various targets (such as vehicle type, quantity, speed, and location) are modeled. Shallower layers are responsible for extracting local structural information such as edges and textures, while deeper layers obtain a comprehensive representation reflecting the overall traffic trend by fusing relationships between regions. Based on this, a feature set that reflects the overall congestion, traffic efficiency, and density distribution of the current traffic environment in the current time period is formed, i.e., the current traffic flow features. Finally, by quantitatively describing the current traffic flow features, such as through classification labels, numerical indicators, or heatmaps, the traffic flow detection results are output.
[0042] Step S103: Based on a preset multimodal coding network, perform data parsing processing on the current image traffic data and historical text traffic data to obtain predicted traffic flow characteristics, and generate traffic flow prediction results for the current traffic environment in future time periods based on the predicted traffic flow characteristics.
[0043] Among them, the multimodal coding network represents a neural network structure for simultaneously receiving and processing data from different data sources (such as current traffic image data and historical traffic image data), and is used to uniformly encode data from different sources in order to extract structured representation features with predictive value. For example, the network can jointly analyze the traffic flow status in the current time period and the traffic flow status trend in the historical time period, thereby establishing the correlation between the traffic flow status trend and specific time or space conditions.
[0044] Among them, the predicted traffic flow characteristics refer to the representative statistical or distributional indicators extracted from the current image traffic data and historical text traffic data, which are used to describe the traffic operation status in the future period. For example, the number of vehicles passing through each road per unit time, average vehicle speed, vehicle density, lane saturation, intersection queue length, etc. in the future period, and the number and density of pedestrians crossing each road, flow direction and trend, etc. in the future period.
[0045] The traffic flow prediction result refers to the structured output information generated by encoding and processing current image traffic data and historical text traffic data, which reflects the trend of traffic operation status changes in a future period. It is used to predict key indicators such as the quantity, distribution, speed or congestion level of traffic flow in the current traffic environment in the future period. For example, the traffic flow prediction result may include the predicted total traffic volume of a road segment in the next 15 minutes, the expected traffic speed of lanes in different directions, or congestion level assessment indicators.
[0046] For example, current image traffic data and historical text traffic data are input into a multimodal coding network. Through the multimodal coding parsing operation of the network, the coupling relationship between the two types of data in the temporal and spatial dimensions is fully explored to transform them into structured information that can be used to characterize traffic operation status. For instance, during the parsing process, the current image traffic data and historical text traffic data are first temporally encoded separately, yielding a series of state sets with evolutionary characteristics at continuous time points. Then, the state sets corresponding to the current image traffic data and historical text traffic data are fused and temporally encoded to capture the dependency structure between the traffic operation status of the current period and the traffic operation status of historical periods. Based on this, a feature set that reflects the overall congestion, traffic efficiency, and density distribution of the current traffic environment in future periods is formed, i.e., predicted traffic flow features. Finally, by quantitatively describing the predicted traffic flow features, such as through classification labels, numerical indicators, or heat maps, the traffic flow prediction result is output.
[0047] The aforementioned traffic flow detection and prediction method based on UAV feedback collaborative optimization firstly achieves comprehensive perception of the current traffic environment by combining current image traffic data collected by image-capturing equipment with historical text traffic data stored in historical time periods. This ensures the real-time nature and data integrity of subsequent traffic flow detection and prediction. Secondly, a multi-level modeling network is used to perform multi-level modeling processing on the current image traffic data, enabling the extraction of structured current traffic flow features and generating traffic flow detection results suitable for quantitative analysis. Thirdly, a multi-modal coding network is used to perform multi-modal coding processing on the current image traffic data and historical text traffic data, enabling the extraction of predicted traffic flow features reflecting the changing trends of traffic operation status and generating traffic flow prediction results suitable for quantitative analysis. Based on this, through a continuous processing chain of data acquisition, traffic flow detection, and traffic flow prediction, quantitative perception and trend reasoning capabilities of the current traffic environment are achieved, improving the accuracy and effectiveness of traffic flow detection and prediction, and providing reliable data support for traffic scheduling and control.
[0048] In one exemplary embodiment, the multi-level modeling network includes a first type of modeling layer in the backbone structure and a second type of modeling layer in the neck structure; as Figure 2 As shown, based on a preset multi-level modeling network, the current image traffic data is parsed and processed to obtain the current traffic flow characteristics, including steps S201 to S202.
[0049] The backbone structure represents the main computation path in the multi-level modeling network. It is used to receive the original current image traffic data and perform layer-by-layer feature extraction to obtain multi-scale feature representations of the current image traffic data from shallow to deep and from local to global at different depth levels. For example, through continuous modeling operations, features such as vehicle outlines, spatial position relationships, and density distribution of traffic flow areas can be extracted from the image.
[0050] The neck structure represents the intermediate fusion path connecting the backbone structure and the network output module. It is used to align, summarize and fuse the feature data of multiple scales output by the backbone structure, enhance the information complementarity between feature data of different scales, and output a unified comprehensive feature representation with stronger modeling capabilities. For example, it can fuse high-resolution vehicle details with low-resolution global traffic flow status to form a unified representation reflecting the overall traffic operation status.
[0051] The first type of modeling layer represents a set of modeling units deployed in the backbone structure to extract image features layer by layer, which are used to extract feature data of different scales from the original image data; the second type of modeling layer represents a set of modeling units deployed in the neck structure to fuse feature data of different scales, which are used to spatially align and integrate the feature data output by the backbone structure, and form the current traffic flow features output uniformly by the neck structure.
[0052] Step S201: Based on the first type of modeling layer, perform data feature extraction processing on the current image traffic data to obtain feature data of different scales output by the backbone structure.
[0053] For example, current image traffic data is input into the backbone structure of a multi-level modeling network. The first type of modeling layer in the backbone structure performs multi-level feature extraction processing on the current image traffic data to extract feature data that reflects different data structure relationship scales. This feature extraction process is not simply a matter of scaling or stacking layers of image data; rather, it involves constructing multiple parallel or serial modeling paths to parse the image data at multiple data organization levels. These modeling paths have different structural depths, feature capture ranges, or channel combinations, enabling modeling of features at different levels and with different data structure relationships within the image data. Through the joint operation of these multiple modeling paths, the backbone structure can extract feature representations corresponding to multiple data structure relationship scales from the current image traffic data. These features differ in their expression level, abstraction level, and correlation ability, reflecting the multiple coupling characteristics between local data structures and global relationships in the traffic scene.
[0054] Step S202: Based on the second type of modeling layer, feature data at different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow features.
[0055] For example, feature data from different data structure relation scales output by the backbone structure are input into the neck structure of a multi-level modeling network. The second type of modeling layer in the neck structure fuses these feature data to integrate them into the current traffic flow features. During data fusion, firstly, the feature data from different modeling paths are structurally coordinated to achieve a unified organizational form, ensuring alignment in dimensions such as spatial coordinates, semantic labels, or channel distribution. Secondly, after coordination, the feature data represented by different data structure relation scales are synthesized. The fusion logic is no longer limited to spatial stacking or channel splicing, but uses weighted calculation, attention mechanisms, or relational mapping to semantically unify the data structure relations expressed in the feature data at different scales. This aims to preserve the independent information carried by the feature data at each scale to the greatest extent possible, while simultaneously achieving collaborative expression of feature data at each scale in a unified representation space. Based on this, the final output of the current traffic flow features is a structured expression formed by fusing multiple data structure relation scales. It possesses the ability to perceive complex structural relationships in traffic scenarios, accurately describes the detailed distribution of local traffic operation states, and includes an abstract representation of the overall traffic operation state.
[0056] In this embodiment, firstly, feature extraction processing is performed on the current image traffic data at different data structure relationship scales according to the first type of modeling layer, thereby constructing multiple sets of feature data reflecting different data structure relationships, improving the ability to express multi-level data structure relationships in complex traffic scenarios; secondly, the second type of modeling layer is used to fuse multiple sets of feature data with scale differences, thereby integrating the data structure relationships of feature data at each layer in a unified feature space to obtain the current traffic flow features, enhancing the perception of traffic operation status in both local and overall aspects; based on this, by constructing and fusing features from different data structure relationship scales, a comprehensive modeling of the current image traffic data is achieved, forming current traffic flow features with hierarchical distribution and unified expressive ability, thus providing a reliable processing basis for generating traffic flow detection results.
[0057] In an exemplary embodiment, the first type of modeling layer includes a spatial relationship modeling layer and a channel relationship modeling layer, and the second type of modeling layer includes a semantic relationship modeling layer; such as Figure 3 As shown, based on the first type of modeling layer, data feature extraction processing is performed on the current image traffic data to obtain feature data of different scales of the backbone structure output, including steps S301 to S303; based on the second type of modeling layer, feature data of different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow features, including step S304.
[0058] Among them, the spatial relationship modeling layer represents a network module used to identify the spatial structural relationships between different pixels in image data, so as to model the relative relationships and distribution characteristics of elements in the image in terms of location, such as modeling the mutual positional relationships between lane boundaries, vehicle gathering areas and road structures in image data.
[0059] Among them, the channel relationship modeling layer represents a network module used to analyze the semantic expression relationship between various channels in image data, so as to model the collaborative or contrastive patterns of different channels in semantic features. For example, modeling the response relationship between the vehicle contour channel and the texture channel in image data can enhance the recognition ability of specific semantic targets. Optionally, a channel represents the information carrier of different dimensions in the feature expression of image data, which is used to record the response data of the image in terms of color, texture, edge or other semantic aspects. For example, in a color image, the three color components of RGB correspond to three channels. In the intermediate layer of a deep neural network, each channel may represent the response intensity distribution to a certain feature (such as vehicle contour, road texture or boundary lines).
[0060] The semantic relationship modeling layer represents a network module used to construct the relationships between semantic objects with semantic meaning in image data, in order to model the interaction relationships between different semantic objects in traffic behavior logic. For example, modeling the response relationship between vehicles and traffic lights or the passage constraints between pedestrians and crosswalks in image data. Furthermore, semantic objects represent regions in image data that have clear semantic labels or functional meanings, which are used to form the basic units of relationship analysis in the semantic modeling layer. For example, vehicles, pedestrians, traffic lights, lane lines and other elements with semantic meaning in traffic image data.
[0061] Step S301: Based on the spatial relationship modeling layer, perform spatial relationship modeling processing on different pixels in the current image traffic data to obtain the first feature data that structurally reflects the spatial relationship.
[0062] The first feature data represents the feature set output by the spatial relationship modeling layer, which is used to express the spatial structural relationship in the image data. It is used to describe the correlation information between pixels or regions in terms of location and structure, such as to reflect the continuity of lane structure or the distribution density of vehicles on the road surface.
[0063] For example, the current image traffic data is input into the spatial relationship modeling layer for spatial relationship modeling processing to establish the spatial distribution relationship between each pixel in the current image traffic data. In this process, each pixel in the current image traffic data is treated as a node with independent attributes, and the spatial distribution relationship between all nodes is modeled. This identifies a set of spatial structure-related features, such as the spatial relative positions, density characteristics, and structural coupling relationships of each element in the current image traffic data, i.e., the first feature data. This first feature data not only preserves the local details of the image data but also establishes the topological relationships between elements on a larger scale, enabling the identification of spatial organization patterns such as densely populated vehicle areas, lane boundaries, and intersection structures.
[0064] Step S302: Based on the channel relationship modeling layer, channel relationship modeling is performed on different channels in the first feature data to obtain the second feature data that structurally reflects the channel relationship.
[0065] The second feature data represents the feature set output by the channel relationship modeling layer, which is used to express the semantic combination relationship between different channels. It is used to enhance the aggregation and expression capabilities of semantic features in image data, such as to strengthen the recognition of the channel combination that expresses the concept of "vehicle" together with the front and taillights.
[0066] For example, the first feature data is input into the channel relationship modeling layer for channel relationship modeling processing to construct semantic connections from multiple channels of feature expression. This involves extracting the most correlated and semantically clear combined features from different channel expression dimensions. In this process, each channel represents the response of a certain type of feature in the image data, such as edges, textures, and color variations. The channel relationship modeling layer analyzes the activation of each channel across the entire image region, capturing the consistency and differences between channels, and mapping this relationship into a more compact and expressive feature set, namely the second feature data. This second feature data enhances the image semantic recognition capability and organizes the expression relationships between channels in a structured manner, enabling an understanding of how the expression dimensions of each channel interact to constitute semantics.
[0067] Step S303: Use the first feature data and the second feature data as feature data of different scales output by the backbone structure.
[0068] Step S304: Based on the semantic relationship modeling layer, perform semantic relationship modeling processing on different semantic objects in the first feature data and the second feature data to obtain the output of the neck structure that reflects the semantic relationship, and integrate the output of the neck structure into the current traffic flow feature.
[0069] For example, the first feature data and the second feature data are used as feature data of different scales output by the backbone structure and input into the semantic relationship modeling layer for semantic relationship modeling processing to extract semantic objects with clear semantic meaning and establish the interaction relationship between various semantic objects. In this process, the semantic relationship modeling layer combines the first feature data extracted by the spatial relationship modeling layer with the second feature data extracted by the channel relationship modeling layer. Through correlation analysis, context construction, and other methods, it identifies the dependency relationship between different semantic objects in traffic behavior logic, such as the dynamic response of vehicles to traffic lights and the crossing behavior of pedestrians to crosswalks, and obtains the output of the neck structure. Finally, the output of the neck structure is integrated into the current traffic flow characteristics, which includes spatial relationships, channel relationships, and semantic relationships, and can support higher-level traffic operation status determination and trend analysis.
[0070] In this embodiment, firstly, spatial relationship modeling is performed on different pixels in the current image traffic data according to the spatial relationship modeling layer to obtain first feature data that structurally reflects spatial relationships; secondly, channel relationship modeling is performed on different channels in the first feature data according to the channel relationship modeling layer to obtain second feature data that structurally reflects channel relationships; thirdly, semantic relationship modeling is performed on different semantic objects in the first and second feature data according to the semantic relationship modeling layer to obtain the output of the neck structure that structurally reflects semantic relationships. Based on this, current traffic flow features containing spatial structure, channel expression, and semantic linkage are generated, which can realize a layer-by-layer modeling process from spatial structure to semantic context and construct a multi-scale fusion feature system suitable for deep analysis of traffic images.
[0071] In one exemplary embodiment, the semantic relationship modeling layer includes a first semantic relationship modeling layer, a second semantic relationship modeling layer, and a third semantic relationship modeling layer; such as Figure 4 As shown, based on the semantic relationship modeling layer, semantic relationship modeling processing is performed on different semantic objects in the first feature data and the second feature data to obtain the output of the neck structure that reflects the semantic relationship, including steps S401 to S404.
[0072] Among them, the first semantic relationship modeling layer, the second semantic relationship modeling layer, and the third semantic relationship modeling layer represent different semantic relationship modeling layers deployed sequentially in the neck structure according to the data flow direction.
[0073] Step S401: Based on the first semantic relationship modeling layer, perform semantic relationship modeling processing on the second feature data to obtain the first intermediate feature data.
[0074] The first intermediate feature data represents the intermediate feature representation obtained by the first semantic relationship modeling layer after performing semantic relationship modeling on the second feature data, and is used to extract the preliminary semantic interaction relationship between semantic objects in the channel dimension.
[0075] Step S402: Based on the second semantic relationship modeling layer, perform semantic relationship modeling processing on the first feature data and the first intermediate feature data to obtain the second intermediate feature data.
[0076] The second intermediate feature data represents the intermediate feature representation obtained by the second semantic relationship modeling layer after jointly modeling the first feature data and the first intermediate feature data. It is used to integrate the interaction characteristics between spatial structure and channel semantics and extract the contextual relationship of cross-scale semantic objects.
[0077] Step S403: Based on the third semantic relationship modeling layer, perform semantic relationship modeling processing on the second feature data and the second intermediate feature data to obtain the third intermediate feature data.
[0078] Among them, the third intermediate feature data is an intermediate feature representation obtained by the third semantic relationship modeling layer after jointly modeling the second feature data and the second intermediate feature data. It is used to further enhance the distribution consistency and expression integrity of channel semantics in the semantic context after fusing spatial structure and channel semantics.
[0079] Step S404: The second intermediate feature data and the third intermediate feature data are used as the output of the neck structure.
[0080] For example, Figure 5 A schematic diagram of a road traffic flow detection algorithm based on real-time UAV data and multi-level relationship modeling is shown. The multi-level modeling network corresponding to this road traffic flow detection algorithm includes Input, Backbone, Neck and Head.
[0081] In the multi-level modeling network, the current image traffic data X acquired in real time by the UAV is input to the Backbone part through the Input part; in the Backbone part, the current image traffic data X is processed by Stage Layer 1 (i.e., the first stage feature extraction layer) to obtain the feature map. ;Will Each pixel is treated as a node, and fine-grained spatial relationship modeling is performed using graph convolutional networks (GCNs) in the SR Layer (i.e., the spatial relationship modeling layer) to obtain a structured representation of spatial relationships. ,in, The calculation method can be referred to formula (1):
[0082] (1)
[0083] in, The channel dimensionality reduction function reduces computation by decreasing the number of channels, and its parameterized representation enhances the algorithm's learnability. Indicates to After function The result obtained by processing Indicates to After the function The transpose of the result obtained after processing.
[0084] If the input is That is to say If the number of channels is c, the height is h, and the width is w, then There are h×w feature points in space, therefore, for Construct a feature map with h×w nodes, and then based on the number of channels c, The shape is adjusted to c×hw, based on equation (1) for the shape c×hw Spatial relationship modeling is performed to obtain Then Restored to its initial shape, resulting in a shape of c×h×w. .
[0085] Will The input is fed into Stage Layer 2 (the second-stage feature extraction layer) for feature extraction. On one hand, the feature extraction result is used as an output of the Backbone part; on the other hand, the feature extraction result is fed into Stage Layer 3 (the third-stage feature extraction layer) for further feature extraction to obtain... .
[0086] On the one hand, As an output of the Backbone section, on the other hand, The input is fed into the CR Layer (i.e., the channel relationship layer), where a graph convolutional network performs fine-grained channel relationship modeling to obtain a structured representation of the channel relationships. ,in, The calculation method can be referred to formula (2):
[0087] (2)
[0088] in, It represents a spatial downsampling function, which reduces computation by reducing spatial dimensions and enhances the learnability of the algorithm through parameterization. Indicates to After the function The result obtained by processing Indicates to After the function The transpose of the result obtained after processing.
[0089] If the input is That is to say If the number of channels is c, the height is h, and the width is w, then... The shape is adjusted to c×hw, based on equation (2) for the shape c×hw Perform channel relationship modeling to obtain Then Restored to its initial shape, resulting in a shape of c×h×w. .
[0090] Will The input is fed into Stage Layer 4 (the fourth stage feature extraction layer) for feature extraction, resulting in... ,Will As an output of the Backbone section.
[0091] For example Figure 5 As shown, The input is fed into Upsample Layer 1 (the first upsampling layer) in the Neck section for upsampling processing. The result of this upsampling process is then compared with the output of Stage Layer 3. The inputs are all fed into ConncatLayer1 (the first connection layer) for concatenation, resulting in... .
[0092] Will The input is fed into Mamba Layer 1 (the first semantic relation modeling layer) for semantic relation modeling processing, resulting in a structured representation that reflects the semantic relations. ;in, The calculation method can be referred to formula (3):
[0093] (3)
[0094] in, Represents a convolution function. Indicates to After the function The result obtained after processing. Regarding the input... By establishing local and global relationships between different feature points in the deeper layers of the network, semantic modeling is achieved, and the output is obtained. .
[0095] Will The input is fed into Upsample Layer 2 (the second upsampling layer) for upsampling processing. The result of this upsampling process is then combined with the feature extraction result from Stage Layer 2 and fed into Conncat Layer 2 (the second connection layer) for concatenation. .
[0096] Will The input is fed into Mamba Layer 2 (the second semantic relation modeling layer) for semantic relation modeling, resulting in a structured representation that reflects the semantic relations. , The calculation method can be referred to formula (3); on the one hand, the calculation method can be referred to formula (3); on the other hand, the calculation method can be referred to formula (3). As an output of the Neck section, on the other hand, The input is fed into Conv Layer1 (the first convolutional layer) for convolution processing, and the result is then fed into Conv Layer2 (the second convolutional layer) for further convolution processing to obtain... .
[0097] Will The upsampling results from Upsample Layer 1 are input together with those from Mamba Layer 3 (the third semantic relation modeling layer) for semantic relation modeling, resulting in a structured representation of semantic relations. , The calculation method can be referred to formula (3), and the calculation method can be used to calculate the result. As an output of the Neck section.
[0098] For example Figure 5 As shown, the output of Mamba Layer2 With Mamba Layer3 output The data are input into the Detection (prediction head) section of the Header for analysis, resulting in traffic flow detection results determined based on the current image traffic data X.
[0099] Understandably, designing a multi-level relationship modeling structure with SR Layer and CR Layer in the Backbone part can more effectively capture long-distance spatial and channel dependencies, improving the accuracy and robustness of feature extraction in the Backbone part. Furthermore, introducing a Mamba Layer in the Neck part, through Mamba's selective mechanism, can effectively handle the relationship modeling between different data blocks while maintaining linear time complexity, alleviating the modeling constraints of convolutional neural networks, and providing advanced modeling capabilities similar to Transformers. This further enhances the relationship modeling of different objects in the network and makes the network focus more on important information while ignoring some interfering information.
[0100] In this embodiment, in the neck structure, firstly, the semantic objects in the second feature data are modeled according to the first semantic relationship modeling layer to extract the first intermediate feature data with a preliminary semantic structure in the channel dimension; secondly, the first feature data and the first intermediate feature data are jointly modeled according to the second semantic relationship modeling layer to strengthen the semantic linkage between spatial structure and channel expression, thereby obtaining the second intermediate feature data; thirdly, the second feature data and the second intermediate feature data are deeply semantically fused according to the third semantic relationship modeling layer to form a high-order semantic expression across scales and channels, thereby obtaining the third intermediate feature data; based on this, according to the complementary semantic structure characteristics of the second intermediate feature data and the third intermediate feature data, they are jointly used as the output of the neck structure, thereby enabling the phased construction of a hierarchical expression structure from local semantic response to global semantic association, improving the modeling capability of multi-objective complex semantic relationships in image data.
[0101] In one exemplary embodiment, the multimodal coding network includes a text coding layer, a visual coding layer, and a text-visual coding layer; as... Figure 6 As shown, based on a preset multimodal coding network, the current image traffic data and historical text traffic data are parsed and processed to obtain predicted traffic flow features, including steps S501 to S504.
[0102] Among them, the text encoding layer represents the encoding structure used to extract text features from historical text traffic data, so as to transform the text descriptions in the historical records into vector representations with contextual semantic information. For example, text content such as "This section of the road was severely congested during the evening rush hour last Friday" is encoded into a feature expression that can participate in subsequent calculations.
[0103] The visual coding layer represents the coding structure used to extract visual features from the current image traffic data, in order to identify visual information such as targets, spatial relationships and traffic flow features in the image data, such as extracting feature vectors with traffic semantics, such as vehicle distribution, lane structure and traffic status in the image data.
[0104] Among them, the text visual coding layer represents a multimodal coding structure used to align and interact between the visual features of current traffic flow and the text features of historical traffic flow. It is used to bridge and fuse data expressions of text and image, so that they can establish a connection in the same semantic space. For example, it can perform correlation analysis between the description of "high traffic density" in the text and the vehicle clustering pattern in the image.
[0105] Step S501: Based on the text encoding layer, perform text feature extraction processing on the historical text traffic data to obtain historical traffic flow text features.
[0106] Among them, the historical traffic flow text feature is a structured vector data output by the text encoding layer, used to represent the text content features in historical traffic data.
[0107] For example, historical traffic data typically records traffic descriptions from a previous period, potentially including weather conditions, road condition changes, traffic congestion levels, construction alerts, or signal adjustments. Based on this, the text encoding layer first serializes the original input text, converting it into a consistent vector representation according to predefined segmentation or embedding rules. Then, it constructs contextual dependencies between words, sentences, or paragraphs through nested encoding structures, transforming the original discrete language information into continuous feature representations usable for subsequent multimodal fusion processing. Furthermore, during encoding, the encoding structure must not only capture semantic dependencies between words but also consider the expression of temporal order and scene relevance. The final generated historical traffic flow text features are a set of numerical vectors with contextual logic, used to express the traffic conditions, trends, and influencing factors reflected in the historical context.
[0108] Step S502: Based on the visual coding layer, perform visual feature extraction processing on the current image traffic data to obtain the current traffic flow visual features.
[0109] Among them, the current traffic flow visual features are structured vector data output by the visual coding layer, which are used to represent the image content features in the current image traffic data.
[0110] For example, in the visual encoding layer, the original input image is first pre-processed and standardized to ensure that the image data adapts to the network structure's input requirements in terms of size, proportion, and color space. Then, a series of computational units analyze the image data layer by layer, extracting visual representation information at different levels from bottom to top. For instance, the initial layer focuses on extracting fine-grained local features such as edges and textures, the middle layer focuses on target shape and distribution, and the deep layer constructs an abstract representation of the overall traffic state. These progressively extracted features are organized into a continuous tensor structure, encompassing vehicle distribution, lane occupancy, road congestion levels, and other traffic-related visual elements in the current image traffic data. The final visual features of the current traffic flow are a set of spatially structured feature maps, preserving the spatial relationships between targets and the overall traffic structure information in the current image traffic data.
[0111] Step S503: Based on the text visual encoding layer, perform data alignment processing on the current traffic flow visual features and historical traffic flow text features at the text data level, visual data level, and text-visual data level respectively to obtain the aligned data corresponding to each data level.
[0112] Step S504: The aligned data corresponding to each data layer are fused to obtain the predicted traffic flow characteristics.
[0113] For example, the historical traffic flow text features output by the text encoding layer and the current traffic flow visual features output by the visual encoding layer are input into the text-visual encoding layer. This coordinates the differences in expression, information dimension, and structural scale between the two modalities. For instance: first, at the text data level, the current traffic flow visual features and historical traffic flow text features are aligned to the text semantic dimension; second, at the visual data level, the current traffic flow visual features and historical traffic flow text features are aligned to the image visual dimension; third, at the text visual data level, the current traffic flow visual features and historical traffic flow text features are aligned to the combined dimension of text semantics and image visuals. Based on this, a mapping relationship between the features of the two modalities is constructed, aligning the "congestion" state mentioned in the text with densely trafficked areas in the image, or mapping the "construction location" described in the text with closed road sections in the image. Finally, the output consists of data alignment results corresponding to each data layer. Each data alignment result maps the specific relationship and mutual reference between text semantics and visual images. The data alignment results of each layer are then integrated to obtain a unified structure for predicting traffic flow features.
[0114] In this embodiment, firstly, text feature extraction processing is performed on historical text traffic data according to the text encoding layer to obtain historical traffic flow text features, thereby extracting a structured textual expression of traffic operation status within a historical time period. Secondly, visual feature extraction processing is performed on current image traffic data according to the visual encoding layer to obtain current traffic flow visual features, thereby extracting a structured visual expression of traffic operation status within the current time period. Thirdly, the current traffic flow visual features and historical traffic flow text features are aligned at multiple data levels according to the text visual encoding layer, thereby establishing a mapping relationship between text semantics and visual images between the two modalities. Fourthly, according to the multi-layer aligned data fusion processing mechanism, the multi-modal features are uniformly integrated into predicted traffic flow features. Based on this, through hierarchical encoding, step-by-step alignment and fusion mechanisms, a unified prediction feature system for heterogeneous traffic data is constructed, improving the completeness and robustness of traffic flow trend modeling.
[0115] In one exemplary embodiment, the multimodal coding network further includes a bidirectional long-short time series modeling layer; such as Figure 7 As shown, based on the text encoding layer, text feature extraction processing is performed on historical text traffic data to obtain historical traffic flow text features, including steps S601 to S602.
[0116] Among them, the bidirectional long and short time series modeling layer represents a time series structure modeling module used to simultaneously model the positive and negative time dependencies in the feature sequence, so as to extract the dynamic relationship and causal connection between the preceding and following semantics in the input sequence from the time dimension. For example, in the traffic text sequence, it can model the sequential logic of "morning rush hour causes congestion" and also identify the negative dependency of "the current smooth flow is due to the previous traffic management measures".
[0117] Step S601: Based on the text encoding layer, perform text feature extraction processing on the historical text traffic data to obtain the initial historical traffic flow text features.
[0118] Step S602: Based on the bidirectional long and short time series modeling layer, the initial historical traffic flow text features are processed by forward time series modeling and reverse time series modeling to obtain historical traffic flow text features.
[0119] For example, Figure 8 This diagram illustrates the structure of a future road traffic flow prediction algorithm based on joint optimization of historical text traffic data and real-time feedback from drones. Its implementation is as follows: Current image traffic data... The data is input into the Vision Encoder (i.e., the visual encoding layer) for visual feature extraction to obtain the current traffic flow visual features; historical text traffic data is then processed. The data is input into BERT (Bidirectional Encoder Representations from Transformers) for text feature extraction to obtain the initial historical traffic flow text features.
[0120] The initial historical traffic flow text features are input into a BiLSTM (Bidirectional Long Short-Term Memory). Multiple LSTMs (Long Short-Term Memory) deployed bidirectionally in the BiLSTM perform forward and reverse time-series modeling on the initial historical traffic flow text features to obtain the historical traffic flow text features.
[0121] The current traffic flow visual features and historical traffic flow text features are input into the text visual data alignment module (i.e., the text visual encoding layer). The text visual data alignment module performs data alignment processing on the current traffic flow visual features and historical traffic flow text features at the text data level, the visual data level, and the text visual data level, respectively, to obtain the aligned data corresponding to each data level.
[0122] The aligned data corresponding to each data layer are fused to obtain the predicted traffic flow features. The predicted traffic flow features are then sequentially processed through multiple Conv Blocks (i.e., convolutional blocks) for layer-by-layer convolution. The final result obtained from the convolution process is input into the prediction head for analysis to obtain the traffic flow prediction result Y.
[0123] For example, the traffic flow prediction result Y can be calculated by referring to equation (4):
[0124] (4)
[0125] in, This indicates visual encoding operations based on Vision Encoder. This represents BERT-based text encoding operations. This represents a bidirectional long-short time series modeling operation based on BiLSTM. This indicates a text-visual bimodal data alignment operation based on the text-visual data alignment module. This represents a convolution operation based on multiple Conv Blocks. Indicates the predictor head operation; symbol " " indicates element-wise multiplication.
[0126] In equation (4), through right Perform visual encoding operations to extract the visual features of the current traffic flow. ;pass right Perform text encoding operations to extract initial historical traffic flow text features. and combined Bidirectional long and short time series modeling operations are performed on the initial historical traffic flow text features to obtain historical traffic flow text features. To establish temporal correlations of text features; through a method based on The established learnable Transformer structure incorporates the visual features of current traffic flow. Historical traffic flow text features Perform text-visual bimodal data alignment and fusion to obtain predicted traffic flow features. ; will predict traffic flow characteristics After sequentially performing convolution and prediction head operations, the traffic prediction regression result for the road several hours in the future is obtained, namely the traffic flow prediction result Y.
[0127] In this embodiment, firstly, text feature extraction processing is performed on historical text traffic data based on the text encoding layer to extract initial historical traffic flow text features with a temporal organizational structure; secondly, forward and reverse time correlation modeling processing is performed on the initial historical traffic flow text features based on the bidirectional long and short time series modeling layer to obtain historical traffic flow text features containing causal relationships and overall semantic integrity; based on this, by constructing a combined mechanism of text semantic extraction and bidirectional time modeling, the expressive power of historical text traffic data in the time dimension is effectively enhanced, providing clear and logically coherent text feature support for subsequent cross-modal fusion and trend prediction.
[0128] In one exemplary embodiment, such as Figure 9 As shown, based on the text visual encoding layer, the current traffic flow visual features and historical traffic flow text features are aligned at the text data level, the visual data level, and the text-visual data level, respectively, to obtain the aligned data corresponding to each data level, including steps S701 to S704.
[0129] Step S701: Based on the historical traffic flow text features, the current traffic flow visual features and the historical traffic flow text features are aligned at the text data level through the text visual encoding layer to obtain the first aligned data at the text data level.
[0130] Among them, the first alignment data represents the data result obtained by semantically aligning the current traffic flow visual features with the historical traffic flow text features as a benchmark at the text data level. It is used to reflect the mapping relationship of visual features in the text semantic organization structure. For example, based on the text description of "congestion on a certain road section", the activation degree of the visual features of the corresponding area in the image is enhanced, and a semantic consistency representation is constructed.
[0131] Step S702: Based on the current traffic flow visual features, the current traffic flow visual features and historical traffic flow text features are aligned at the visual data level through a text visual encoding layer to obtain the second aligned data at the visual data level.
[0132] The second alignment data refers to the data results obtained at the visual data level after spatially aligning the text features of historical traffic flow with the current traffic flow visual features as a benchmark. It is used to express the relocation process of text semantics in the image spatial structure. For example, based on the structural features of dense traffic flow areas in the image, the expression "slow traffic" in the text is relocated and embedded.
[0133] Step S703: Based on the splicing features corresponding to the historical traffic flow text features and the current traffic flow visual features, the splicing features are aligned at the text visual data level through the text visual encoding layer to obtain the third aligned data at the text visual data level.
[0134] Among them, the third alignment data refers to the data results obtained by uniformly modeling and aligning the splicing features of historical traffic flow text features and current traffic flow visual features at the text visual data level. This data is used to integrate the commonalities and differences between the two modalities in semantics and structure, and to construct a cross-modal joint semantic space. For example, the time description in the text and the spatial state in the image are encoded in parallel to form a multimodal composite expression of traffic trends in a specific scenario.
[0135] Step S704: The first alignment data, the second alignment data, and the third alignment data are taken as the aligned data corresponding to each data layer.
[0136] The aligned data refers to the multi-layer data alignment output result composed of the first aligned data, the second aligned data, and the third aligned data, which is used to provide unified input features with consistent structure and clear semantics for subsequent feature fusion and traffic flow prediction tasks.
[0137] For example, Figure 10 A schematic diagram of a data alignment algorithm based on a text-visual data alignment module is shown. Its implementation is as follows: At the text data level, the visual features of the current traffic flow are respectively... Historical traffic flow text features Layer normalization (LN) is performed, which will be... The obtained layer normalization results are consistent with those obtained from... The obtained layer normalization result is calculated through element-wise multiplication to obtain the first product result. Then, the first product result is transformed using the Tanh (Hyperbolic Tangent Function) to obtain the first transformed result. Finally, the result is... The obtained layer normalization result and the first transformation result are calculated by element-wise multiplication to obtain the second product result. The first alignment data at the text data level is obtained by performing element-wise addition on the result of the second product. .
[0138] At the visual data level, historical traffic flow text features are respectively... Visual characteristics of current traffic flow Perform layer normalization processing, which will be... The obtained layer normalization results are consistent with those obtained from... The obtained layer normalization result is calculated using element-wise multiplication to obtain the third product result, which is then subjected to a Tanh transform to obtain the second transform result; then the result is further processed by... The obtained layer normalization result and the second transformation result are calculated through element-wise multiplication to obtain the fourth product result. The second alignment data at the visual data level is obtained by performing element-wise addition on the result of the fourth product. .
[0139] At the textual visual data level, historical traffic flow text features Visual characteristics of current traffic flow The input is fed into ConCat (the connection layer) for splicing, resulting in... and The corresponding splicing features are then processed by layer normalization. The layer normalization result is then calculated using element-wise multiplication to obtain the fifth product. This fifth product is then subjected to Tanh transformation to obtain the third transformation result. The layer normalization result and the third transformation result are then calculated using element-wise multiplication to obtain the sixth product. Finally, the splicing features and the sixth product are calculated using element-wise addition to obtain the third alignment data at the text visual data level. .
[0140] Then align the first data Second alignment data Third alignment data The data is input to Fusion (the fusion operation module) for fusion processing. This involves integrating feature-aligned data from multiple data layers through methods such as splicing, weighting, or attention mechanisms to obtain the predicted traffic flow features output by the text-visual data alignment module. .
[0141] For example, the first alignment data The calculation method can be referred to in equation (5), the second alignment data The calculation method can be referred to in equation (6), third alignment data The calculation method can be referred to in equation (7) to predict traffic flow characteristics. The calculation method can be referred to formula (8):
[0142] (5)
[0143] (6)
[0144] (7)
[0145] (8)
[0146] In equations (5), (6), and (7), and ( This indicates that it is implemented through a single fully connected layer. Figure 9 The layer normalization operation of LN is shown; ( This indicates that it is achieved through a pre-defined normalization function. Figure 9 The transformation operation of Tanh is shown; [·,·] indicates that it is implemented through the ConCat function. Figure 9 The ConCat feature concatenation operation shown is performed by dimension, for example in equation (7). Indicates to and Perform feature splicing operation; symbol " " indicates element-wise multiplication.
[0147] In equation (8), This represents the feature fusion operation at different levels, in order to achieve... Figure 9 The Fusion shown will , , The fusion process is performed to obtain aligned enhanced features that reflect various different characteristics, i.e., traffic flow prediction features. .
[0148] Understandably, the text-visual data alignment module can represent a novel bimodal data adaptive alignment network model designed based on a multi-head self-attention module. By fully considering the characteristics of both text and visual data, it can, on the one hand, achieve deep data fusion, fully explore the complementary information between different modalities, and thus powerfully enhance feature representation, providing a richer and more accurate feature foundation for subsequent analysis; on the other hand, it can significantly improve model performance and generalization ability. By introducing diverse data, it improves the model's modeling capabilities, allowing the model to exhibit excellent adaptability and stability even in complex and ever-changing new data environments.
[0149] In this embodiment, firstly, based on historical traffic flow text features, data alignment processing is performed on the current traffic flow visual features and historical traffic flow text features at the text data level, thereby achieving semantic adaptation of visual information. Secondly, based on the current traffic flow visual features, data alignment processing is performed on the current traffic flow visual features and historical traffic flow text features at the visual data level, thereby achieving the positioning and mapping of text semantics in the spatial structure. Thirdly, based on the splicing features of the current traffic flow visual features and historical traffic flow text features, joint alignment processing is performed, thereby constructing a unified representation with cross-modal semantic consistency. Based on this, by constructing a multi-layered mechanism of text-dominated, visual-dominated, and joint alignment, a structurally complete and semantically coordinated multi-modal alignment feature system is formed, providing a stable and consistent representation foundation for subsequent fusion and prediction.
[0150] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0151] Based on the same inventive concept, this application also provides a traffic flow detection and prediction device based on UAV feedback collaborative optimization for implementing the aforementioned traffic flow detection and prediction method based on UAV feedback collaborative optimization. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the traffic flow detection and prediction device based on UAV feedback collaborative optimization provided below can be found in the limitations of the traffic flow detection and prediction method based on UAV feedback collaborative optimization described above, and will not be repeated here.
[0152] In one exemplary embodiment, such as Figure 11 As shown, a traffic flow detection and prediction device based on UAV feedback collaborative optimization is provided, including: an acquisition module 101, a detection module 102, and a prediction module 103, wherein:
[0153] The acquisition module 101 is used to acquire current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in the historical time period in the current traffic environment. The image capturing device is multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment.
[0154] The detection module 102 is used to perform data parsing and processing on the current image traffic data based on a preset multi-level modeling network to obtain the current traffic flow characteristics, and generate the traffic flow detection result of the current traffic environment in the current time period based on the current traffic flow characteristics.
[0155] The prediction module 103 is used to perform data parsing and processing on the current image traffic data and historical text traffic data based on a preset multimodal coding network to obtain predicted traffic flow characteristics, and generate traffic flow prediction results for the current traffic environment in the future time period based on the predicted traffic flow characteristics.
[0156] In an exemplary embodiment, the detection module 102 is further configured to: perform data feature extraction processing on the current image traffic data based on the first type of modeling layer to obtain feature data of different scales of the backbone structure output; perform fusion processing on the feature data of different scales based on the second type of modeling layer to obtain the output of the neck structure, and integrate the output of the neck structure into the current traffic flow features.
[0157] In an exemplary embodiment, the detection module 102 is further configured to: perform spatial relationship modeling processing on different pixels in the current image traffic data based on the spatial relationship modeling layer to obtain first feature data that structurally reflects spatial relationships; perform channel relationship modeling processing on different channels in the first feature data based on the channel relationship modeling layer to obtain second feature data that structurally reflects channel relationships; output feature data of different scales as the backbone structure using the first feature data and the second feature data as the backbone structure; perform semantic relationship modeling processing on different semantic objects in the first feature data and the second feature data based on the semantic relationship modeling layer to obtain the output of the neck structure that structurally reflects semantic relationships, and integrate the output of the neck structure into the current traffic flow features.
[0158] In an exemplary embodiment, the detection module 102 is further configured to: perform semantic relationship modeling processing on the second feature data based on the first semantic relationship modeling layer to obtain first intermediate feature data; perform semantic relationship modeling processing on the first feature data and the first intermediate feature data based on the second semantic relationship modeling layer to obtain second intermediate feature data; perform semantic relationship modeling processing on the second feature data and the second intermediate feature data based on the third semantic relationship modeling layer to obtain third intermediate feature data; and use the second intermediate feature data and the third intermediate feature data as the output of the neck structure.
[0159] In an exemplary embodiment, the prediction module 103 is further configured to: perform text feature extraction processing on historical text traffic data based on the text encoding layer to obtain historical traffic flow text features; perform visual feature extraction processing on current image traffic data based on the visual encoding layer to obtain current traffic flow visual features; perform data alignment processing on the current traffic flow visual features and historical traffic flow text features at the text data level, visual data level, and text-visual data level, respectively, based on the text-visual encoding layer to obtain aligned data corresponding to each data level; and fuse the aligned data corresponding to each data level to obtain predicted traffic flow features.
[0160] In an exemplary embodiment, the prediction module 103 is further configured to: perform text feature extraction processing on historical text traffic data based on the text encoding layer to obtain initial historical traffic flow text features; and perform forward time series modeling processing and reverse time series modeling processing on the initial historical traffic flow text features based on the bidirectional long and short time series modeling layer to obtain historical traffic flow text features.
[0161] In an exemplary embodiment, the prediction module 103 is further configured to: align the current traffic flow visual features with the historical traffic flow text features at the text data level using a text visual encoding layer, based on historical traffic flow text features, to obtain first aligned data at the text data level; align the current traffic flow visual features with the historical traffic flow text features at the visual data level using a text visual encoding layer, based on the current traffic flow visual features, to obtain second aligned data at the visual data level; align the spliced features corresponding to the historical traffic flow text features and the current traffic flow visual features at the text visual data level using a text visual encoding layer, based on historical traffic flow text features and the spliced features, to obtain third aligned data at the text visual data level; and use the first aligned data, the second aligned data, and the third aligned data as the aligned data corresponding to each data level.
[0162] The modules in the aforementioned traffic flow detection and prediction device based on UAV feedback collaborative optimization can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0163] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above embodiments.
[0164] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above embodiments.
[0165] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0166] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0167] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A traffic flow detection and prediction method based on UAV feedback collaborative optimization, characterized in that, The method includes: In the current traffic environment, current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in historical time periods are acquired. The image capturing device consists of multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment. Based on a preset multi-level modeling network, the current image traffic data is parsed and processed to obtain the current traffic flow characteristics. Based on the current traffic flow characteristics, the traffic flow detection result of the current traffic environment in the current time period is generated. Based on a preset multimodal coding network, the current image traffic data and the historical text traffic data are parsed and processed to obtain predicted traffic flow features. Based on the predicted traffic flow features, a traffic flow prediction result for the current traffic environment in the future time period is generated. The multimodal coding network includes a text coding layer, a visual coding layer, and a text-visual coding layer. The preset multimodal coding network performs data parsing processing on the current image traffic data and the historical text traffic data to obtain predicted traffic flow features, including: Based on the text encoding layer, text feature extraction processing is performed on the historical text traffic data to obtain historical traffic flow text features; based on the visual encoding layer, visual feature extraction processing is performed on the current image traffic data to obtain current traffic flow visual features; based on the text visual encoding layer, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the text data level, visual data level, and text-visual data level, respectively, to obtain aligned data corresponding to each data level; the aligned data corresponding to each data level are fused to obtain predicted traffic flow features. Specifically, based on the text visual encoding layer, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the text data level, visual data level, and text-visual data level, respectively, to obtain the aligned data corresponding to each data level, including: Based on the historical traffic flow text features, the current traffic flow visual features and the historical traffic flow text features are aligned at the text data level through the text visual encoding layer to obtain first aligned data at the text data level; based on the current traffic flow visual features, the current traffic flow visual features and the historical traffic flow text features are aligned at the visual data level through the text visual encoding layer to obtain second aligned data at the visual data level; based on the splicing features corresponding to the historical traffic flow text features and the current traffic flow visual features, the splicing features are aligned at the text visual data level through the text visual encoding layer to obtain third aligned data at the text visual data level; the first aligned data, the second aligned data, and the third aligned data are used as the aligned data corresponding to each data level.
2. The method according to claim 1, characterized in that, The multi-level modeling network includes a first type of modeling layer in the backbone structure and a second type of modeling layer in the neck structure. The pre-defined multi-level modeling network performs data parsing processing on the current image traffic data to obtain current traffic flow characteristics, including: Based on the first type of modeling layer, data feature extraction processing is performed on the current image traffic data to obtain feature data of different scales output by the backbone structure. Based on the second type of modeling layer, feature data at different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow features.
3. The method according to claim 2, characterized in that, The first type of modeling layer includes a spatial relationship modeling layer and a channel relationship modeling layer, and the second type of modeling layer includes a semantic relationship modeling layer; Based on the first type of modeling layer, the current image traffic data is processed to extract features, resulting in feature data of different scales output by the backbone structure, including: Based on the spatial relationship modeling layer, spatial relationship modeling processing is performed on different pixels in the current image traffic data to obtain first feature data that structurally reflects spatial relationships; Based on the channel relationship modeling layer, channel relationship modeling processing is performed on different channels in the first feature data to obtain second feature data that structurally reflects the channel relationship; The first feature data and the second feature data are used as feature data of different scales output by the backbone structure; Based on the second type of modeling layer, feature data at different scales are fused to obtain the output of the neck structure. The output of the neck structure is then integrated into the current traffic flow features, including: Based on the semantic relationship modeling layer, semantic relationship modeling is performed on different semantic objects in the first feature data and the second feature data to obtain the output of the neck structure that reflects the semantic relationship. The output of the neck structure is then integrated into the current traffic flow features.
4. The method according to claim 3, characterized in that, The semantic relation modeling layer includes a first semantic relation modeling layer, a second semantic relation modeling layer, and a third semantic relation modeling layer; The semantic relationship modeling layer performs semantic relationship modeling processing on different semantic objects in the first feature data and the second feature data to obtain a structured output of the neck structure that reflects the semantic relationship, including: Based on the first semantic relationship modeling layer, semantic relationship modeling processing is performed on the second feature data to obtain the first intermediate feature data; Based on the second semantic relationship modeling layer, semantic relationship modeling processing is performed on the first feature data and the first intermediate feature data to obtain the second intermediate feature data. Based on the third semantic relationship modeling layer, semantic relationship modeling processing is performed on the second feature data and the second intermediate feature data to obtain the third intermediate feature data. The second intermediate feature data and the third intermediate feature data are used as the output of the neck structure.
5. The method according to claim 1, characterized in that, The multimodal coding network also includes a bidirectional long-short time series modeling layer; The step of extracting text features from the historical text traffic data based on the text encoding layer to obtain historical traffic flow text features includes: Based on the text encoding layer, text feature extraction processing is performed on the historical text traffic data to obtain the initial historical traffic flow text features; Based on the bidirectional long and short time series modeling layer, the initial historical traffic flow text features are processed by forward time series modeling and reverse time series modeling to obtain historical traffic flow text features.
6. A traffic flow detection and prediction device based on UAV feedback collaborative optimization, characterized in that, The device includes: The acquisition module is used to acquire current image traffic data collected by a preset image capturing device in the current time period and historical text traffic data stored in the historical time period in the current traffic environment. The image capturing device is a combination of multiple UAV devices deployed in the current traffic environment to jointly observe the current traffic environment. The detection module is used to perform data parsing processing on the current image traffic data based on a preset multi-level modeling network to obtain the current traffic flow characteristics, and generate the traffic flow detection result of the current traffic environment in the current time period based on the current traffic flow characteristics. The prediction module is used to perform data parsing processing on the current image traffic data and the historical text traffic data based on a preset multimodal coding network to obtain predicted traffic flow features, and generate traffic flow prediction results for the current traffic environment in future time periods based on the predicted traffic flow features. The multimodal coding network includes a text coding layer, a visual coding layer, and a text-visual coding layer. The prediction module is further configured to: extract text features from the historical text traffic data based on the text coding layer to obtain historical traffic flow text features; extract visual features from the current image traffic data based on the visual coding layer to obtain current traffic flow visual features; align the current traffic flow visual features and the historical traffic flow text features at the text data level, the visual data level, and the text-visual data level, respectively, based on the text-visual coding layer to obtain aligned data corresponding to each data level; and fuse the aligned data corresponding to each data level to obtain predicted traffic flow features. The prediction module is further configured to: using the historical traffic flow text features as a reference, perform data alignment processing on the current traffic flow visual features and the historical traffic flow text features at the text data level through the text visual encoding layer to obtain first aligned data at the text data level; using the current traffic flow visual features as a reference, perform data alignment processing on the current traffic flow visual features and the historical traffic flow text features at the visual data level through the text visual encoding layer to obtain second aligned data at the visual data level; using the splicing features corresponding to the historical traffic flow text features and the current traffic flow visual features as a reference, perform data alignment processing on the splicing features at the text visual data level through the text visual encoding layer to obtain third aligned data at the text visual data level; and use the first aligned data, the second aligned data, and the third aligned data as the aligned data corresponding to each data level.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Internet unmanned aerial vehicle capable of achieving continuous endurance
CN105652886A
Data processing method and device, electronic equipment and storage medium
CN115115913A
Pedestrian flow prediction system and method
CN117592599A