Traffic flow detection and prediction method based on unmanned aerial vehicle feedback collaborative optimization

Through multi-level modeling and multi-modal coding network for collaborative optimization of UAVs, traffic data is processed, and the adaptability and accuracy of traffic flow detection and prediction in complex scenarios in the prior art is solved, and more efficient traffic flow detection and prediction is achieved.

CN120449100AActive Publication Date: 2025-08-08ALADDIN UAV (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510596297.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-08
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

In the prior art, the method of collecting image data through a fixed camera for traffic flow detection and prediction is poor in complex traffic scenarios and insufficient prediction accuracy.

Method used

Using a method based on drone feedback collaborative optimization, image and text traffic data are processed through multi-level modeling networks and multi-modal encoding networks, structured features are extracted and traffic flow detection and prediction results are generated.

Benefits of technology

It improves the accuracy and effectiveness of traffic flow detection and prediction, and provides reliable data support for traffic scheduling and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449100A_ABST
    Figure CN120449100A_ABST
Patent Text Reader

Abstract

The invention relates to a traffic flow detection and prediction method based on unmanned aerial vehicle feedback collaborative optimization. The method comprises the following steps: in a current traffic environment, obtaining current image traffic data collected by a preset image shooting device in a current time period and historical text traffic data stored in a historical time period; based on a preset multi-level modeling network, performing data analysis processing on the current image traffic data to obtain current traffic flow characteristics, and generating a traffic flow detection result of the current traffic environment in the current time period according to the current traffic flow characteristics; and based on a preset multi-mode coding network, performing data analysis processing on the current image traffic data and the historical text traffic data to obtain predicted traffic flow characteristics, and generating a traffic flow prediction result of the current traffic environment in the future period according to the predicted traffic flow characteristics. By adopting the method, the accuracy and effectiveness of detecting and predicting the traffic flow can be improved, and reliable data support is provided for traffic scheduling and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of traffic flow detection, and in particular to a traffic flow detection and prediction method based on UAV feedback collaborative optimization. Background Art

[0002] In the field of traffic flow optimization technology, it involves detecting and predicting traffic flow, so as to recommend traffic driving strategies to vehicle drivers.

[0003] In related technologies, fixed cameras are deployed to collect image data, and traditional image processing algorithms are used to estimate indicators such as traffic density and travel speed, thereby achieving identification, detection and trend prediction of the current traffic operation status. However, such methods have the disadvantages of limited ability to express data features and insufficient description of time series evolution laws, resulting in poor adaptability of detection and prediction results to complex traffic scenarios and insufficient prediction accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a traffic flow detection and prediction method, device, computer equipment and computer-readable storage medium based on drone feedback collaborative optimization to address the above technical problems, so as to improve the accuracy and effectiveness of traffic flow detection and prediction, and provide reliable data support for traffic scheduling and control.

[0005] In the first aspect, the present application provides a traffic flow detection and prediction method based on drone feedback collaborative optimization, comprising: In a current traffic environment, current image traffic data collected by a preset image capturing device in a current time period and historical text traffic data stored in a historical time period are obtained, wherein the image capturing device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to perform joint observation of the current traffic environment; Based on a preset multi-level modeling network, the current image traffic data is subjected to data analysis processing to obtain current traffic flow characteristics, and a traffic flow detection result of the current traffic environment in the current time period is generated according to the current traffic flow characteristics; Based on a preset multimodal encoding network, data analysis and processing are performed on the current image traffic data and the historical text traffic data to obtain predicted traffic flow characteristics, and traffic flow prediction results for the current traffic environment in future time periods are generated based on the predicted traffic flow characteristics.

[0006] Secondly, the present application also provides a traffic flow detection and prediction device based on drone feedback collaborative optimization, comprising: An acquisition module is configured to acquire, in a current traffic environment, current image traffic data collected by a preset image capture device during a current period and historical text traffic data stored during historical periods, wherein the image capture device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to perform joint observation of the current traffic environment; a detection module configured to perform data analysis on the current image traffic data based on a preset multi-level modeling network to obtain current traffic flow characteristics, and generate a traffic flow detection result of the current traffic environment in the current time period based on the current traffic flow characteristics; The prediction module is used to perform data analysis and processing on the current image traffic data and the historical text traffic data based on a preset multimodal encoding network to obtain predicted traffic flow characteristics, and generate traffic flow prediction results for the current traffic environment in the future time period based on the predicted traffic flow characteristics.

[0007] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above steps when executing the computer program.

[0008] In a fourth aspect, the present application also provides a computer-readable storage medium on which a computer program is stored, and the computer program implements the above steps when executed by a processor.

[0009] The above-mentioned traffic flow detection and prediction method, device, computer equipment and computer-readable storage medium based on drone feedback collaborative optimization firstly, comprehensively perceive the current traffic environment through the current image traffic data collected by the image capture device in the current time period and the historical text traffic data stored in the historical time period, thereby ensuring the real-time and data integrity of the subsequent detection and prediction of traffic flow; secondly, multi-level modeling processing is performed on the current image traffic data according to the multi-level modeling network, so as to extract structured current traffic flow characteristics and generate traffic flow detection results that can be used for quantitative analysis; thirdly, multi-modal encoding processing is performed on the current image traffic data and the historical text traffic data according to the multi-modal encoding network, so as to extract predicted traffic flow characteristics reflecting the changing trend of traffic operation status and generate traffic flow prediction results that can be used for quantitative analysis; based on this, through the continuous processing chain of data collection, traffic flow detection and traffic flow prediction, quantitative perception and trend reasoning capabilities of the current traffic environment are realized, the accuracy and effectiveness of traffic flow detection and prediction are improved, and reliable data support is provided for traffic scheduling and control. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 1. A flow chart of a traffic flow detection and prediction method based on UAV feedback collaborative optimization in one embodiment; Figure 2 A schematic diagram of a process for generating current traffic flow characteristics based on a multi-level modeling network in one embodiment; Figure 3 A schematic diagram of a process for generating current traffic flow characteristics based on a multi-level modeling network in another embodiment; Figure 4 A schematic diagram of a process for generating neck structure output based on a semantic relationship modeling layer in a multi-level modeling network in one embodiment; Figure 5 A schematic diagram of the structure of a road traffic flow detection algorithm based on real-time drone data and multi-level relationship modeling in one embodiment; Figure 6 A schematic diagram of a process for generating predicted traffic flow features based on a multimodal encoding network in one embodiment; Figure 7 A schematic diagram of a process for generating historical traffic flow text features by combining a text encoding layer and a bidirectional long-short time series modeling layer in a multimodal encoding network in one embodiment; Figure 8 A schematic diagram of the structure of a future road traffic flow prediction algorithm based on joint optimization of historical text traffic data and real-time feedback from drones in one embodiment; Figure 9 A schematic diagram of a process for generating aligned data corresponding to each data level based on a text visual coding layer in a multimodal coding network in one embodiment; Figure 10 A schematic diagram of the structure of a data alignment algorithm based on a text-visual data alignment module in one embodiment; Figure 11 The present invention is a structural block diagram of a traffic flow detection and prediction device based on UAV feedback collaborative optimization in one embodiment. DETAILED DESCRIPTION

[0012] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0013] In one embodiment, Figure 1 As shown, a traffic flow detection and prediction method based on drone feedback collaborative optimization is provided. This embodiment uses the method applied to a server as an example. It is understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S101 to S103.

[0014] Step S101: In the current traffic environment, current image traffic data collected by a preset image capture device in the current time period and historical text traffic data stored in the historical time period are obtained. The image capture device is a plurality of drone devices deployed in the current traffic environment to perform joint observation of the current traffic environment.

[0015] The current traffic environment refers to the traffic operation status within a specified area, that is, it is used to describe the comprehensive traffic conditions such as vehicles, pedestrians, road facilities, and traffic light control duration within the specified area.

[0016] Among them, the image capture device refers to the equipment deployed in the current traffic environment for obtaining real-time traffic images, that is, it is used to obtain traffic data in the form of image data related to factors such as vehicles, pedestrians, and road facilities; the image capture device can refer to a tethered drone satellite hovering in the air, which is equipped with a high-pixel camera and can achieve centimeter-level real-time resolution of ground objects.

[0017] Among them, the current image traffic data refers to the image traffic data collected by the image shooting equipment in the current time period and reflecting the traffic operation status of the current time period; the historical text traffic data refers to the text traffic data stored in text format and reflecting the traffic operation status of the historical time period.

[0018] For example, in a city traffic map, the physical space corresponding to a specific map area can be designated as the current traffic environment. For example, the physical space corresponding to the entire city in the city traffic map can be designated as the current traffic environment, or the physical space corresponding to multiple interconnected main roads in the city traffic map can be designated as the current traffic environment. Furthermore, over time, the network model or algorithm used for traffic flow detection and prediction in this embodiment will generate a large amount of historical data. This historical data reflects the traffic conditions of various roads in the current traffic environment at different historical moments. This historical data is stored in a database in text form and serves as historical text traffic data for the historical period when detecting and predicting traffic flow in the current period.

[0019] The image capture equipment represents tethered drones deployed on a city-by-city basis. This is equivalent to evenly deploying a number of tethered drones over a city, with the camera units in these drones continuously capturing centimeter-level image data in real time, thereby achieving comprehensive visual coverage and global traffic monitoring of the entire urban traffic area. The number, distribution, and spatial distances between tethered drones deployed over the city can be determined by combining factors such as the city's road distribution, city boundary outline, and area size, along with the tethered drones' field of view. This allows the entire city's traffic area to be fully covered through the global field of view formed by these tethered drones. The field of view of a tethered drone can be defined as a circular area defined by the tethered drone's location as the center and the maximum visible distance as the radius.

[0020] Furthermore, the structural composition and corresponding functions of the above-mentioned tethered drone can refer to a continuous-endurance Internet drone in the relevant technology (publication number CN105652886A), which is a tethered drone satellite equipped with a high-pixel camera and a field of view of up to 5 square kilometers; on the one hand, the drone can achieve continuous endurance based on continuous power supply, thereby maintaining a continuous working state; on the other hand, based on higher wind and rain resistance, the drone can continue to work under general weather conditions, based on this, it can achieve real-time and uninterrupted acquisition of centimeter-level precision image data.

[0021] Step S102 , based on a preset multi-level modeling network, the current image traffic data is analyzed and processed to obtain current traffic flow characteristics, and a traffic flow detection result of the current traffic environment in the current period is generated according to the current traffic flow characteristics.

[0022] Among them, the multi-level modeling network representation represents a neural network structure used to extract and analyze traffic flow features at different levels from the original traffic data layer by layer. It is used to convert basic visual elements (such as edges, colors, and motion) in the original image data into structured expression features that can reflect the traffic operation status. For example, the shallow network is used to extract vehicle contours and boundary information, and the deep network is used to infer the overall traffic density or traffic congestion trends in local areas.

[0023] Among them, the current traffic flow characteristics represent representative statistical or distribution indicators extracted from the current image traffic data, which are used to describe the traffic operation status of the current time period, such as the distance between each vehicle and the corresponding terminal in the current time period, the number of vehicles passing through each road per unit time, the average speed, vehicle density, lane saturation, intersection queue length, etc., and the number and density of pedestrians traveling on each road in the current time period, flow direction and trend, etc.

[0024] Among them, the traffic flow detection result represents the structured output information obtained after modeling and processing of the current image traffic data, reflecting the traffic operation status in the current time period. It is used to describe key indicators such as the number, distribution, speed or congestion level of the current traffic environment in the current time period. For example, the traffic flow detection result may include the number of vehicles passing through a certain intersection per unit time in the current time period, the average speed of each lane, and whether there are abnormal traffic incidents in the road section.

[0025] For example, current image traffic data is input into a multi-level modeling network. Through the layer-by-layer feature parsing operation of the multi-level modeling network, the low-level visual features in the original image data are gradually converted into structured information that can be used to characterize traffic operation status. For example, during the parsing process, elements such as target outlines, motion trajectories, and spatial distribution are first extracted from the original image data. Then, various types of targets (such as vehicle type, number, speed, and location) are modeled. The shallower layers are responsible for extracting local structural information such as edges and textures, while the deeper layers integrate the relationships between regions to obtain a comprehensive representation reflecting the overall traffic trend. Based on this, a feature set is formed that can reflect the overall congestion situation, traffic efficiency, and density distribution of the current traffic environment during the current time period, namely the current traffic flow characteristics. Finally, the current traffic flow characteristics are quantitatively described, such as through classification labels, numerical indicators, or heat maps, and the traffic flow detection results are output.

[0026] Step S103: Based on a preset multimodal coding network, the current image traffic data and the historical text traffic data are analyzed and processed to obtain predicted traffic flow characteristics, and a traffic flow prediction result for the current traffic environment in the future period is generated based on the predicted traffic flow characteristics.

[0027] Among them, the multimodal encoding network represents a neural network structure for simultaneously receiving and processing data from different data sources (such as current traffic image data and historical traffic image data), and is used to uniformly encode data from different sources to extract structured expression features with predictive value. For example, the network can jointly analyze the traffic flow status in the current time period and the traffic flow status trend in the historical time period, thereby establishing a correlation between the traffic flow status trend and specific time or spatial conditions.

[0028] Among them, the predicted traffic flow features represent representative statistical or distribution indicators extracted from current image traffic data and historical text traffic data, which are used to describe the traffic operation status in the future time period, such as the number of vehicles passing through each road per unit time in the future time period, the average speed, vehicle density, lane saturation, intersection queue length, etc., and the number and density of pedestrians traveling on each road in the future time period, as well as the flow direction and trend.

[0029] Among them, the traffic flow prediction result represents the structured output information generated after encoding the current image traffic data and historical text traffic data, reflecting the changing trend of traffic operation status in a certain period of time in the future. It is used to predict key indicators such as the number, distribution, speed or congestion level of the current traffic environment in the future period. For example, the traffic flow prediction result may include the predicted value of the total traffic volume of a certain section of road in the next 15 minutes, the expected travel speed of lanes in different directions, or congestion level assessment indicators.

[0030] For example, current image traffic data and historical text traffic data are fed into a multimodal encoding network. Through the network's multimodal encoding parsing operations, the coupling relationship between the two types of data in the temporal and spatial dimensions is fully exploited to convert them into structured information that can be used to characterize traffic operation status. For example, during the parsing process, the current image traffic data and historical text traffic data are first temporally encoded, resulting in a series of state sets with evolutionary characteristics at consecutive time points. Furthermore, the state sets corresponding to the current image traffic data and historical text traffic data are fused and temporally encoded to capture the dependency structure between the traffic operation status of the current time period and that of the historical time periods. Based on this, a feature set is formed that reflects the overall congestion, traffic efficiency, and density distribution of the current traffic environment in the future time period, namely, the predicted traffic flow characteristics. Finally, the predicted traffic flow characteristics are quantitatively described, for example, through classification labels, numerical indicators, or heat maps, and the traffic flow prediction results are output.

[0031] In the above-mentioned traffic flow detection and prediction method based on drone feedback collaborative optimization, first, the current traffic environment is fully perceived through the current image traffic data collected by the image capture device in the current period and the historical text traffic data stored in the historical period, thereby ensuring the real-time and data integrity of the subsequent traffic flow detection and prediction; secondly, the current image traffic data is multi-level modeled according to the multi-level modeling network, so that structured current traffic flow features can be extracted and traffic flow detection results that can be used for quantitative analysis are generated; thirdly, the current image traffic data and historical text traffic data are multi-modally encoded according to the multi-modal coding network, so that predicted traffic flow features reflecting the changing trend of traffic operation status can be extracted and traffic flow prediction results that can be used for quantitative analysis are generated; based on this, through the continuous processing chain of data collection, traffic flow detection, and traffic flow prediction, the quantitative perception and trend reasoning capabilities of the current traffic environment are realized, the accuracy and effectiveness of traffic flow detection and prediction are improved, and reliable data support is provided for traffic scheduling and control.

[0032] In an exemplary embodiment, the multi-level modeling network includes a first type of modeling layer in the backbone structure and a second type of modeling layer in the neck structure; Figure 2 As shown, based on the preset multi-level modeling network, the current image traffic data is analyzed and processed to obtain the current traffic flow characteristics, including steps S201 to S202.

[0033] Among them, the backbone structure represents the main computing path in the multi-level modeling network, which is used to receive the original current image traffic data and perform layer-by-layer feature extraction to obtain multi-scale feature representations from shallow to deep and from local to global in the current image traffic data at different depth levels. For example, through continuous modeling operations, features such as vehicle contours, spatial position relationships, and density distribution of traffic areas are extracted from the image.

[0034] Among them, the neck structure represents the intermediate fusion path connecting the backbone structure and the network output module. It is used to align, aggregate and fuse the multiple scale feature data output by the backbone structure, enhance the information complementarity between feature data of different scales, and output a unified comprehensive feature representation with stronger modeling capabilities. For example, high-resolution vehicle details are integrated with low-resolution global traffic flow status to form a unified representation reflecting the entire traffic operation status.

[0035] Among them, the first type of modeling layer represents a group of modeling units deployed in the backbone structure for extracting image features layer by layer, and is used to extract feature data of different scales from the original image data; the second type of modeling layer represents a group of modeling units deployed in the neck structure for fusing feature data of different scales, and is used to perform spatial alignment and information integration on the feature data output by the backbone structure, and form the current traffic flow characteristics uniformly output by the neck structure.

[0036] In step S201 , based on the first type of modeling layer, data feature extraction processing is performed on the current image traffic data to obtain feature data of different scales output by the backbone structure.

[0037] Exemplarily, the current image traffic data is input into the backbone structure of the multi-level modeling network, and the first type of modeling layer in the backbone structure performs multi-level feature extraction processing on the current image traffic data to extract feature data that can reflect different data structure relationship scales; wherein, the feature extraction process is not just a simple resizing or hierarchical stacking of the image data, but rather a process of parsing the image data at multiple data organization levels by constructing multiple parallel or serial modeling paths. These modeling paths have different structural depths, feature capture ranges or channel combination methods, so that features at different levels and different data structure relationships in the image data can be modeled separately; through the joint operation of multiple modeling paths, the backbone structure can extract feature representations corresponding to multiple data structure relationship scales from the current image traffic data. These features differ in expression level, degree of abstraction, association ability, etc., and respectively reflect the multiple coupling characteristics between local data structures and global relationships in traffic scenes.

[0038] In step S202 , based on the second modeling layer, feature data of different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow feature.

[0039] For example, feature data from different data structure relationship scales output by the backbone structure are input into the neck structure of the multi-level modeling network. The second modeling layer in the neck structure fuses these feature data to form the current traffic flow characteristics. During the data fusion process, the feature data from different modeling paths are first structurally coordinated to give them a unified organizational form, ensuring that they can be aligned in dimensions such as spatial coordinates, semantic labels, or channel distribution. After the coordination is completed, the feature data represented by different data structure relationship scales are synthesized. The fusion logic is no longer limited to spatial stacking or channel splicing. Instead, through weighted calculation, attention mechanism, or relational mapping, the data structure relationships expressed in the feature data of different scales are semantically unified. This aims to maximize the preservation of the independent information carried by the feature data at each scale, while simultaneously achieving the coordinated expression of the feature data at each scale in a unified representation space. Based on this, the final output of the current traffic flow characteristics is a structured expression result formed by the fusion of multiple data structure relationship scales. It has the ability to perceive the complex structural relationships in the traffic scene, accurately describe the detailed distribution of local traffic operation status, and also include an abstract representation of the overall traffic operation status.

[0040] In this embodiment, first, according to the first type of modeling layer, feature extraction processing is performed on the current image traffic data at different data structure relationship scales, so that multiple groups of feature data sets reflecting different data structure relationships can be constructed, thereby improving the expression ability of multi-level data structure relationships in complex traffic scenes; secondly, according to the second type of modeling layer, multiple groups of feature data with scale differences are fused, thereby integrating the data structure relationships of feature data of each layer in a unified feature space to obtain the current traffic flow characteristics, thereby enhancing the perception ability of the traffic operation status in locality and overallness; based on this, by constructing and fusing features from different data structure relationship scales, comprehensive modeling of the current image traffic data is achieved, forming current traffic flow characteristics with hierarchical distribution and unified expression ability, thereby providing a reliable processing basis for generating traffic flow detection results.

[0041] In an exemplary embodiment, the first type of modeling layer includes a spatial relationship modeling layer and a channel relationship modeling layer, and the second type of modeling layer includes a semantic relationship modeling layer; Figure 3 As shown, based on the first type of modeling layer, data feature extraction processing is performed on the current image traffic data to obtain feature data of different scales output by the backbone structure, including steps S301 to S303; based on the second type of modeling layer, feature data of different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow feature, including step S304.

[0042] Among them, the spatial relationship modeling layer represents a network module used to identify the spatial structural relationship between different pixels in image data, so as to model the relative positional relationship and distribution characteristics of elements in the image. For example, it can model the relative positional relationship between lane boundaries, vehicle gathering areas, and road structures in image data.

[0043] Among them, the channel relationship modeling layer represents a network module used to analyze the semantic expression relationship between each channel in the image data, so as to model the collaborative or contrasting patterns of different channels in semantic features. For example, the response relationship between the vehicle contour channel and the texture channel is modeled in the image data to enhance the recognition ability of specific semantic targets; optionally, the channel represents the information carrier of different dimensions in the feature expression of the image data, which is used to record the response data of the image in color, texture, edge or other semantic aspects respectively. For example, in a color image, the three color components RGB correspond to three channels. In the middle layer of the deep neural network, each channel may represent the response intensity distribution to a certain feature (such as vehicle contour, road texture or boundary line).

[0044] Among them, the semantic relationship modeling layer represents a network module used to construct the mutual relationship between semantic objects with semantic meaning in image data, so as to model the interactive relationship between different semantic objects in the traffic behavior logic, such as modeling the response relationship between vehicles and traffic lights or the traffic constraints between pedestrians and sidewalks in image data; furthermore, semantic objects represent areas with clear semantic labels or functional meanings in image data, which are used to constitute the basic units of relationship analysis in the semantic modeling layer, such as vehicles, pedestrians, traffic lights, lane lines and other elements with semantic meanings in traffic image data.

[0045] Step S301 : Based on the spatial relationship modeling layer, spatial relationship modeling processing is performed on different pixels in the current image traffic data to obtain first feature data that structurally reflects the spatial relationship.

[0046] Among them, the first feature data represents the feature set output by the spatial relationship modeling layer, which is used to express the spatial structural relationship in the image data, and is used to describe the correlation information between pixels or regions in position and structure, such as used to reflect the continuity of lane structure or the distribution density of vehicles on the road surface.

[0047] Exemplarily, the current image traffic data is input into the spatial relationship modeling layer for spatial relationship modeling processing to establish the spatial distribution relationship between each pixel in the current image traffic data. During this process, each pixel in the current image traffic data is treated as a node with independent attributes, and the spatial distribution relationship between all nodes is modeled. This identifies a set of spatial structure-related features, such as the spatial position, density characteristics, and structural coupling relationships of each element in the current image traffic data, namely the first feature data. This first feature data not only retains the local details of the image data but also establishes the topological relationships between each element on a larger scale, enabling the identification of spatial organizational forms such as densely populated areas, lane boundaries, and intersection structures.

[0048] Step S302 : Based on the channel relationship modeling layer, channel relationship modeling processing is performed on different channels in the first feature data to obtain second feature data that structurally reflects the channel relationship.

[0049] Among them, the second feature data represents the feature set output by the channel relationship modeling layer, which is used to express the semantic combination relationship between different channels, and is used to enhance the aggregation and expression capabilities of semantic features in image data, for example, to enhance the recognition of the channel combination of the front and taillights that jointly express the concept of "vehicle".

[0050] Exemplarily, the first feature data is input into the channel relationship modeling layer for channel relationship modeling processing to construct semantic connections from multiple channels of feature expression, that is, to extract the most relevant and semantically clear combination features between different channel expression dimensions. In this process, each channel represents the response of a certain type of feature in the image data, such as edges, textures, color changes, etc. The channel relationship modeling layer analyzes the activation of each channel in the entire image area, captures the consistency and difference between channels, and maps this relationship into a more compact and expressive feature set, namely the second feature data. This second feature data strengthens the recognition ability of image semantics and organizes the expression relationship between channels in a structural way, making it possible to understand how the expression dimensions of each channel interact to form semantics.

[0051] Step S303: The first feature data and the second feature data are used as feature data of different scales output by the backbone structure.

[0052] Step S304: Based on the semantic relationship modeling layer, semantic relationship modeling is performed on different semantic objects in the first feature data and the second feature data to obtain a structured output of the neck structure that reflects the semantic relationship, and the output of the neck structure is integrated into the current traffic flow feature.

[0053] Exemplarily, the first feature data and the second feature data are used as feature data of different scales output by the backbone structure and input into the semantic relationship modeling layer for semantic relationship modeling processing to extract semantic objects with clear semantic meanings and establish interactive relationships between the various semantic objects. In this process, the semantic relationship modeling layer combines the first feature data extracted by the spatial relationship modeling layer with the second feature data extracted by the channel relationship modeling layer. Through correlation analysis, context construction, etc., the dependency relationship between different semantic objects in traffic behavior logic is identified, such as the dynamic response of vehicles and traffic lights, the travel behavior of pedestrians and sidewalks, etc., to obtain the output of the neck structure. Ultimately, the output of the neck structure is integrated into the current traffic flow characteristics, which include spatial relationships, channel relationships, and semantic relationships, and can support higher-level traffic operation status judgment and trend analysis.

[0054] In this embodiment, first, spatial relationship modeling is performed on different pixels in the current image traffic data according to the spatial relationship modeling layer to obtain first feature data that structuredly reflects the spatial relationship; second, channel relationship modeling is performed on different channels in the first feature data according to the channel relationship modeling layer to obtain second feature data that structuredly reflects the channel relationship; third, semantic relationship modeling is performed on different semantic objects in the first feature data and the second feature data according to the semantic relationship modeling layer to obtain an output of a neck structure that structuredly reflects the semantic relationship. Based on this, the current traffic flow characteristics including spatial structure, channel expression and semantic linkage are generated, which can realize a layer-by-layer modeling process from spatial structure to semantic context, and construct a multi-scale fusion feature system suitable for deep analysis of traffic images.

[0055] In an exemplary embodiment, the semantic relationship modeling layer includes a first semantic relationship modeling layer, a second semantic relationship modeling layer, and a third semantic relationship modeling layer; Figure 4 As shown, based on the semantic relationship modeling layer, semantic relationship modeling processing is performed on different semantic objects in the first feature data and the second feature data to obtain a structured output of the neck structure that reflects the semantic relationship, including steps S401 to S404.

[0056] The first semantic relationship modeling layer, the second semantic relationship modeling layer, and the third semantic relationship modeling layer respectively represent different semantic relationship modeling layers deployed in sequence according to the data flow direction in the neck structure.

[0057] Step S401 : Based on the first semantic relationship modeling layer, semantic relationship modeling is performed on the second feature data to obtain first intermediate feature data.

[0058] Among them, the first intermediate feature data represents the intermediate feature representation obtained after the first semantic relationship modeling layer performs semantic relationship modeling on the second feature data, and is used to extract the preliminary semantic interaction relationship between semantic objects in the channel dimension.

[0059] Step S402 : Based on the second semantic relationship modeling layer, semantic relationship modeling is performed on the first feature data and the first intermediate feature data to obtain second intermediate feature data.

[0060] Among them, the second intermediate feature data represents the intermediate feature representation obtained by jointly modeling the first feature data and the first intermediate feature data by the second semantic relationship modeling layer, which is used to integrate the interaction characteristics between spatial structure and channel semantics and refine the contextual relationship of cross-scale semantic objects.

[0061] Step S403 : Based on the third semantic relationship modeling layer, semantic relationship modeling is performed on the second feature data and the second intermediate feature data to obtain third intermediate feature data.

[0062] Among them, the third intermediate feature data represents the intermediate feature representation obtained by jointly modeling the second feature data and the second intermediate feature data by the third semantic relationship modeling layer, and is used to further enhance the distribution consistency and expression integrity of the channel semantics in the semantic context after the fusion of spatial structure and channel semantics.

[0063] Step S404: Use the second intermediate feature data and the third intermediate feature data as output of the neck structure.

[0064] For example, Figure 5 The figure shows a structural diagram of a road traffic flow detection algorithm based on real-time drone data and multi-level relationship modeling. The multi-level modeling network corresponding to the road traffic flow detection algorithm includes Input (input layer), Backbone (i.e., backbone structure), Neck (i.e., neck structure) and Header (i.e., head structure).

[0065] In the multi-level modeling network, the current image traffic data X acquired by the drone in real time is input to the Backbone part through the Input part; in the Backbone part, the feature extraction of the current image traffic data X is performed through Stage Layer 1 (i.e., the first stage feature extraction layer) to obtain the feature map ;Will Each pixel is taken as a node, and fine-grained spatial relationship modeling is performed through the graph convolutional network (GCN) in the SR Layer (i.e., spatial relationship modeling layer) to obtain a structured spatial relationship. ,in, The calculation method can refer to formula (1): (1) in, Represents the channel dimensionality reduction function, which reduces the amount of computation by reducing the number of channels and enhances the learnability of the algorithm through parameterized representation; Express Through the function The result obtained by processing, Express Through the function Perform the transpose operation on the processed result.

[0066] If the input , which means The number of channels is c, the height is h, and the width is w, then There are h×w feature points in the space, so for Construct a feature map with h×w nodes, and then transform The shape of is adjusted to c×hw, based on formula (1) Perform spatial relationship modeling to obtain , and then Restore to the original shape and get a shape of c×h×w .

[0067] Will The feature extraction is input to Stage Layer 2 (i.e., the second stage feature extraction layer) for feature extraction. On the one hand, the feature extraction result is used as an output of the Backbone part. On the other hand, the feature extraction result is input to Stage Layer 3 (i.e., the third stage feature extraction layer) for further feature extraction to obtain .

[0068] On the one hand, As an output of the Backbone part, on the other hand, Input to CR Layer (i.e. channel relationship layer), and the graph convolution network in CR Layer performs fine-grained channel relationship modeling to obtain a structured channel relationship. ,in, The calculation method can refer to formula (2): (2) in, Represents a spatial downsampling function, which reduces the amount of computation by reducing the spatial dimension and enhances the learnability of the algorithm through parameterized representation; Express Through the function The result obtained by processing, Express Through the function Perform the transpose operation on the processed result.

[0069] If the input , which means The number of channels is c, the height is h, and the width is w, then The shape of is adjusted to c×hw, based on formula (2) Perform channel relationship modeling to obtain , and then Restore to the original shape and get a shape of c×h×w .

[0070] Will Input to Stage Layer4 (i.e. the fourth stage feature extraction layer) for feature extraction, and obtain ,Will As an output of the Backbone part.

[0071] For example Figure 5 As shown, The input is sent to Upsample Layer1 (the first upsampling layer) of Neck for upsampling, and the result of upsampling is compared with the output of Stage Layer3. Input them together to ConncatLayer1 (the first connection layer) for splicing and get .

[0072] Will Input to Mamba Layer 1 (i.e. the first semantic relationship modeling layer) for semantic relationship modeling processing to obtain a structured representation of the semantic relationship. ;in, The calculation method can refer to formula (3): (3) in, represents a convolution function, Express Through the function The result of processing. By establishing local and global relationships between different feature points in the deep layer of the network, semantic modeling is achieved and the output is obtained. .

[0073] Will The result of upsampling is input to Upsample Layer2 (i.e. the second upsampling layer) for upsampling, and then the result of upsampling and the feature extraction result output by Stage Layer2 are input to Conncat Layer2 (i.e. the second connection layer) for splicing. .

[0074] Will Input to Mamba Layer2 (the second semantic relationship modeling layer) for semantic relationship modeling processing to obtain a structured representation of the semantic relationship. , The calculation method of can refer to formula (3); on the one hand, As an output of the Neck part, on the other hand, The input is sent to Conv Layer 1 (i.e. the first convolutional layer) for convolution processing, and the result of the convolution processing is then sent to Conv Layer 2 (i.e. the second convolutional layer) for further convolution processing to obtain .

[0075] Will The upsampling result output by Upsample Layer1 is input to Mamba Layer3 (the third semantic relationship modeling layer) for semantic relationship modeling, and a structured representation of the semantic relationship is obtained. , The calculation method of can refer to formula (3), and As an output of the Neck part.

[0076] For example Figure 5 As shown, the output of Mamba Layer2 Output from Mamba Layer3 The Detection (i.e., prediction head) is input into the Header part for analysis to obtain the traffic flow detection result determined based on the current image traffic data X.

[0077] It can be understood that the multi-level relationship modeling structure of SR Layer and CR Layer designed in the Backbone part can more effectively capture long-distance spatial and channel dependencies, improving the accuracy and robustness of feature extraction in the Backbone part; furthermore, the introduction of Mamba Layer in the Neck part, through Mamba's selective mechanism, effectively handles the relationship modeling between different data blocks while maintaining linear time complexity, alleviates the modeling constraints of convolutional neural networks, and provides advanced modeling capabilities similar to Transformers, thereby further enhancing the relationship modeling of different objects in the network and enabling the network to pay more attention to some important information and ignore some interference information.

[0078] In this embodiment, in the neck structure, first, the semantic objects in the second feature data are modeled according to the first semantic relationship modeling layer, so as to extract the first intermediate feature data with preliminary semantic structure under the channel dimension; secondly, the first feature data and the first intermediate feature data are jointly modeled according to the second semantic relationship modeling layer, so as to strengthen the semantic linkage relationship between the spatial structure and the channel expression, and obtain the second intermediate feature data; thirdly, the second feature data and the second intermediate feature data are deep semantic fusion modeled according to the third semantic relationship modeling layer, so as to form a high-order semantic expression across scales and channels, and obtain the third intermediate feature data; based on this, according to the complementary characteristics of the semantic structures of the second intermediate feature data and the third intermediate feature data, they are combined as the output of the neck structure, so that a hierarchical expression structure from local semantic response to global semantic association can be constructed in stages, thereby improving the modeling ability of multi-target complex semantic relationships in image data.

[0079] In an exemplary embodiment, the multimodal encoding network includes a text encoding layer, a visual encoding layer, and a text-visual encoding layer; Figure 6 As shown, based on a preset multimodal encoding network, data parsing and processing are performed on the current image traffic data and the historical text traffic data to obtain predicted traffic flow characteristics, including steps S501 to S504.

[0080] Among them, the text encoding layer represents the encoding structure used to extract text features from historical text traffic data, so as to convert the text descriptions in the historical records into vector representations with contextual semantic information. For example, the text content such as "The road section was severely congested during the evening rush hour last Friday" is encoded into a feature expression that can participate in subsequent calculations.

[0081] Among them, the visual coding layer represents the coding structure used to extract visual features of the current image traffic data to identify visual information such as targets, spatial relationships and flow characteristics in the image data, for example, extracting feature vectors with traffic semantics such as vehicle distribution, lane structure and traffic status in the image data.

[0082] Among them, the text visual encoding layer represents a multimodal encoding structure used to align and interactively model the visual features of current traffic flow and the text features of historical traffic flow. It is used to bridge and fuse the data expressions of text and image modes so that they can establish connections in the same semantic space. For example, the "high traffic density" described in the text can be associated with the vehicle aggregation pattern in the image for analysis.

[0083] Step S501 : Based on the text coding layer, text feature extraction processing is performed on the historical text traffic data to obtain the historical traffic flow text features.

[0084] The historical traffic flow text features represent structured vector data output by the text encoding layer and used to represent the text content features in the historical text traffic data.

[0085] For example, historical text traffic data usually records traffic descriptions from a previous period, which may include information such as weather conditions, road condition changes, traffic congestion levels, construction reminders, or signal adjustments. Based on this, the text encoding layer first serializes the original input text, that is, converts the original input text into a vector expression with consistent form according to preset word segmentation or embedding rules. Then, through a nested encoding structure, the contextual dependencies between words, sentences, or paragraphs are constructed, thereby converting the original discrete language information into a continuous feature representation that can be used for subsequent multimodal fusion processing. Furthermore, during the encoding process, the encoding structure must not only capture the semantic dependencies between word meanings, but also consider the expression of temporal order and scene relevance. The final generated historical traffic flow text features are a set of numerical vectors with contextual logic, which are used to express the traffic status, trends, and influencing factors reflected in the historical context.

[0086] Step S502: Based on the visual coding layer, visual feature extraction processing is performed on the current image traffic data to obtain the current traffic flow visual features.

[0087] The current traffic flow visual feature represents the structured vector data output by the visual coding layer and used to represent the image content features in the current image traffic data.

[0088] For example, in the visual encoding layer, the original input image is first preliminarily regularized and preprocessed to ensure that the image data adapts to the input requirements of the network structure in terms of size, scale, and color space. Secondly, the image data is analyzed layer by layer through a series of computing units, and visual expression information at different levels is extracted from the bottom up. For example, the initial layer focuses on extracting fine-grained local features such as edges and textures, the middle layer focuses on the shape and distribution of targets, and the deep layer constructs an abstract expression of the overall traffic status. The features gradually extracted at these levels are organized into a continuous tensor structure, covering the vehicle distribution, lane occupancy, road congestion level, and other visual elements related to the traffic status in the current image traffic data. The final visual features of the current traffic flow are a set of spatially structured feature maps that retain the spatial relationship between targets in the current image traffic data and the overall traffic structure information.

[0089] In step S503, based on the text visual coding layer, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the text data level, the visual data level, and the text visual data level to obtain aligned data corresponding to each data level.

[0090] In step S504, the aligned data corresponding to each data level are fused to obtain predicted traffic flow characteristics.

[0091] Exemplarily, the historical traffic flow text features output by the text encoding layer and the current traffic flow visual features output by the visual encoding layer are input into the text-visual encoding layer together, and the differences in the expression methods, information dimensions, and structural scales of the features of the two modalities are coordinated. For example: first, at the text data level, the current traffic flow visual features and the historical traffic flow text features are guided to the dimension of text semantics for data alignment processing; secondly, at the visual data level, the current traffic flow visual features and the historical traffic flow text features are guided to the dimension of image vision for data alignment processing; thirdly, at the text-visual data level, the current traffic flow visual features and the historical traffic flow text features are guided to the combined dimension of text semantics and image vision for data alignment processing; based on this, a mapping relationship between the features of the two modalities is constructed, and the "congestion" state mentioned in the text is aligned with the vehicle-dense area in the image, or the "construction location" described in the text is mapped with the closed road section in the image. Finally, the output is the data alignment results corresponding to each data level. Each layer of data alignment results maps the specific association and mutual reference between text semantics and visual images. The final feature integration of each layer of data alignment results is then performed to obtain the predicted traffic flow characteristics presented in a unified structure.

[0092] In this embodiment, first, text feature extraction is performed on historical text traffic data according to the text coding layer to obtain historical traffic flow text features, thereby extracting a structured text expression of the traffic operation status in the historical period; secondly, visual feature extraction is performed on current image traffic data according to the visual coding layer to obtain current traffic flow visual features, thereby extracting a structured visual expression of the traffic operation status in the current period; thirdly, the current traffic flow visual features and historical traffic flow text features are aligned at multiple data levels according to the text visual coding layer, thereby establishing a mapping relationship between text semantics and visual images between the two modalities; thirdly, according to the fusion processing mechanism of multi-layer aligned data, the multimodal features are unified and integrated into predicted traffic flow features. Based on this, through hierarchical coding, step-by-step alignment and fusion mechanism, a unified prediction feature system for heterogeneous traffic data is constructed, which improves the integrity and robustness of traffic flow trend modeling.

[0093] In an exemplary embodiment, the multimodal encoding network further includes a bidirectional long-short temporal sequence modeling layer; Figure 7 As shown, based on the text coding layer, text feature extraction processing is performed on the historical text traffic data to obtain historical traffic flow text features, including steps S601 to S602.

[0094] The bidirectional long-short time series modeling layer represents a temporal structure modeling module used to simultaneously model the forward and reverse temporal dependencies in feature sequences, thereby extracting the dynamic relationship and causal connection between the previous and subsequent semantics in the input sequence from the temporal dimension. For example, in a traffic text sequence, it can model the sequential logic of "the morning rush hour causes congestion" and also identify the reverse dependency of "the current unimpeded state is due to the previous diversion measures."

[0095] Step S601 : Based on the text coding layer, text feature extraction processing is performed on the historical text traffic data to obtain initial historical traffic flow text features.

[0096] Step S602 : Based on the bidirectional long-short time series modeling layer, the initial historical traffic flow text features are subjected to forward time series modeling processing and reverse time series modeling processing to obtain historical traffic flow text features.

[0097] For example, Figure 8 The schematic diagram of the structure of the future road traffic flow prediction algorithm based on the joint optimization of historical text traffic data and real-time feedback from drones is shown. The implementation method is as follows: the current image traffic data Input to Vision Encoder (i.e., visual encoding layer) for visual feature extraction and processing to obtain the current traffic flow visual features; historical text traffic data The data is input into BERT (Bidirectional Encoder Representations from Transformers) for text feature extraction to obtain the initial historical traffic flow text features.

[0098] The initial historical traffic flow text features are input into the BiLSTM (i.e., bidirectional long short-term time series modeling layer, Bidirectional Long Short-Term Memory). Multiple LSTMs (i.e., long short-term time series modeling layer, Long Short-Term Memory) bidirectionally deployed in the BiLSTM perform forward and reverse time series modeling on the initial historical traffic flow text features to obtain the historical traffic flow text features.

[0099] The current traffic flow visual features and the historical traffic flow text features are input into the text visual data alignment module (i.e., the text visual encoding layer). The text visual data alignment module performs data alignment processing on the current traffic flow visual features and the historical traffic flow text features at the text data level, visual data level, and text visual data level, respectively, to obtain the aligned data corresponding to each data level.

[0100] The aligned data corresponding to each data layer are fused to obtain the predicted traffic flow characteristics. The predicted traffic flow characteristics are then sequentially convolved layer by layer through multiple Conv Blocks (i.e., convolution blocks). The final result obtained from the convolution processing is input into the prediction head for analysis to obtain the traffic flow prediction result Y.

[0101] For example, the traffic flow prediction result Y can be calculated by referring to formula (4): (4) in, Represents the visual encoding operation based on Vision Encoder, Represents the BERT-based text encoding operation, Represents the bidirectional long-short time series modeling operation based on BiLSTM, Represents the text-vision bimodal data alignment operation based on the text-vision data alignment module. Represents a convolution operation based on multiple Conv Blocks, Indicates the prediction head operation; the symbol " ” means element-wise multiplication.

[0102] In formula (4), by right Perform visual encoding operations to extract the current traffic flow visual features ;pass right Perform text encoding operations to extract the initial historical traffic flow text features , and combined with Perform bidirectional long-short time series modeling on the initial historical traffic flow text features to obtain historical traffic flow text features , in order to establish the temporal association relationship of text features; The established learnable Transformer structure transforms the current traffic flow visual features and historical traffic flow text features Perform text-visual bimodal data alignment and fusion to obtain predicted traffic flow characteristics ; Traffic flow characteristics will be predicted After the convolution operation and the prediction head operation, the traffic prediction regression result for the road in the next few hours is obtained, that is, the traffic flow prediction result Y.

[0103] In this embodiment, first, text feature extraction is performed on historical text traffic data according to the text encoding layer, so as to extract the initial historical traffic flow text features with a temporal organizational structure; secondly, forward and reverse time correlation modeling is performed on the initial historical traffic flow text features according to the bidirectional long-short time series modeling layer, so as to obtain historical traffic flow text features containing causal correlation and overall semantic integrity; based on this, by constructing a combined mechanism of text semantic extraction and bidirectional time modeling, the expression ability of historical text traffic data in the time dimension is effectively enhanced, providing clear structure and logically coherent text feature support for subsequent cross-modal fusion and trend prediction.

[0104] In an exemplary embodiment, Figure 9 As shown, based on the text visual coding layer, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the text data level, the visual data level and the text visual data level respectively to obtain the aligned data corresponding to each data level, including steps S701 to S704.

[0105] Step S701, based on the historical traffic flow text features, the current traffic flow visual features and the historical traffic flow text features are aligned at the text data level through the text visual coding layer to obtain the first aligned data at the text data level.

[0106] Among them, the first alignment data represents the data results obtained after semantic alignment of the current traffic flow visual features based on the historical traffic flow text features at the text data level. It is used to reflect the mapping relationship of visual features in the text semantic organization structure. For example, based on the "congestion on a certain road section" information described in the text, the activation degree of the visual features of the corresponding area in the image is enhanced, and a semantic consistency representation is constructed.

[0107] Step S702, based on the current traffic flow visual features, the current traffic flow visual features and the historical traffic flow text features are aligned at the visual data level through the text visual coding layer to obtain second aligned data at the visual data level.

[0108] Among them, the second alignment data represents the data results obtained after spatial structural alignment of historical traffic flow text features based on the current traffic flow visual features at the visual data level. It is used to express the repositioning process of text semantics in the image spatial structure. For example, based on the structural features of the dense traffic flow area in the image, the position association and embedding reconstruction of the expression "slow traffic" in the text are performed.

[0109] In step S703, based on the splicing features corresponding to the historical traffic flow text features and the current traffic flow visual features, the splicing features are aligned at the text visual data level through the text visual coding layer to obtain third aligned data at the text visual data level.

[0110] Among them, the third alignment data represents the data results obtained after unified modeling and alignment of the splicing features of historical traffic flow text features and current traffic flow visual features at the text-visual data level. It is used to fuse the commonalities and differences between the two modalities in semantics and structure, and construct a cross-modal joint semantic space. For example, the time description in the text and the spatial state in the image are encoded in parallel to form a multimodal composite expression of traffic trends in specific scenarios.

[0111] In step S704 , the first aligned data, the second aligned data, and the third aligned data are used as aligned data corresponding to each data layer.

[0112] Among them, the aligned data represents the multi-layer data alignment output result composed of the first aligned data, the second aligned data and the third aligned data, which is used to provide unified input features with consistent structure and clear semantics for subsequent feature fusion and traffic flow prediction tasks.

[0113] For example, Figure 10 The structure diagram of a data alignment algorithm based on text-visual data alignment module is shown. The implementation method is as follows: at the text data level, the current traffic flow visual features are respectively and historical traffic flow text features Perform layer normalization (LN, Layer Normalization) processing, which will be The layer normalization results obtained are the same as those obtained by The obtained layer normalization result is calculated by element-by-element multiplication operation to obtain the first product result, and then the first product result is transformed by Tanh (Hyperbolic Tangent Function) to obtain the first transformation result; The obtained layer normalization result and the first transformation result are calculated by element-by-element multiplication operation to obtain the second product result. The first alignment data at the text data level is calculated by element-by-element addition operation with the second product result .

[0114] At the visual data level, the historical traffic flow text features are respectively Visual characteristics of current traffic flow Perform layer normalization, which will be The layer normalization results obtained are the same as those obtained by The obtained layer normalization result is calculated by element-by-element multiplication operation to obtain the third product result, and then the third product result is subjected to Tanh transformation to obtain the second transformation result; The obtained layer normalization result and the second transformation result are calculated by element-by-element multiplication operation to obtain the fourth product result. The second alignment data at the visual data level is calculated by element-by-element addition operation with the fourth product result .

[0115] At the text visual data level, the historical traffic flow text features Visual characteristics of current traffic flow Input to ConCat (i.e. connection layer) for splicing processing, and get and The corresponding splicing features; the splicing features are layer-normalized, and the layer-normalized results obtained by the splicing features are calculated by element-by-element multiplication self-interaction operation to obtain the fifth product result, and then the fifth product result is Tanh transformed to obtain the third transformation result; the layer-normalized results obtained by the splicing features and the third transformation result are calculated by element-by-element multiplication operation to obtain the sixth product result, and the splicing features and the sixth product result are calculated by element-by-element addition operation to obtain the third alignment data at the text visual data level. .

[0116] Then the first alignment data , second alignment data , third alignment data Input to Fusion (i.e. fusion operation module) for fusion processing, such as integrating feature alignment data on multiple data levels through splicing, weighting or attention mechanism, so as to obtain the predicted traffic flow features output by the text-visual data alignment module .

[0117] For example, the first alignment data The calculation method can refer to formula (5), the second alignment data The calculation method can refer to formula (6), the third alignment data The calculation method of can refer to formula (7), predicting traffic flow characteristics The calculation method can refer to formula (8): (5) (6) (7) (8) In formula (5), formula (6), and formula (7), and ( ) means that they are realized by a single-layer fully connected layer Figure 9 The layer normalization operation of the LN shown; ( ) indicates that the normalization function is preset. Figure 9 The Tanh transformation operation shown; [·,·] indicates that it is implemented by the ConCat function Figure 9 The ConCat shown in the figure performs feature concatenation by dimension, for example, in Eq. (7) Express and Perform feature splicing operation; symbol " ” means element-wise multiplication.

[0118] In formula (8), Represents the fusion operation of features at different levels to achieve Figure 9 The Fusion shown will 、 、 Fusion is performed to obtain aligned enhanced features that can reflect various feature characteristics, that is, predicted traffic flow features .

[0119] It can be understood that the text-visual data alignment module can represent a new type of bimodal data adaptive alignment network model designed based on the multi-head self-attention module. On the basis of fully considering the characteristics of text data and visual data, on the one hand, it can realize the deep fusion of data, fully explore the complementary information between each modality, and thus effectively enhance the feature expression, providing a richer and more accurate feature basis for subsequent analysis; on the other hand, it can significantly improve the model performance and generalization ability, and improve the model modeling ability by introducing diversified data, so that the model can also show excellent adaptability and stability in the complex and changing new data environment.

[0120] In this embodiment, first, based on the text features of historical traffic flow, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the text data level, so as to achieve expression adaptation of visual information in the semantic dimension; secondly, based on the current traffic flow visual features, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the visual data level, so as to achieve positioning and mapping of text semantics in the spatial structure; thirdly, based on the splicing features of the current traffic flow visual features and the historical traffic flow text features, joint alignment processing is performed, so as to construct a unified representation of cross-modal semantic consistency; based on this, by constructing a multi-layer mechanism of text-dominant, vision-dominant and joint alignment, a multi-modal alignment feature system with complete structure and semantic coordination is formed, which provides a stable and consistent representation basis for subsequent fusion and prediction.

[0121] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0122] Based on the same inventive concept, the embodiments of the present application also provide a traffic flow detection and prediction device based on drone feedback collaborative optimization for implementing the above-mentioned traffic flow detection and prediction method based on drone feedback collaborative optimization. The implementation solution provided by this device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations of one or more traffic flow detection and prediction device embodiments based on drone feedback collaborative optimization provided below can be found in the above-mentioned limitations of the traffic flow detection and prediction method based on drone feedback collaborative optimization, and will not be repeated here.

[0123] In an exemplary embodiment, Figure 11 As shown, a traffic flow detection and prediction device based on UAV feedback collaborative optimization is provided, including: an acquisition module 101, a detection module 102 and a prediction module 103, wherein: An acquisition module 101 is configured to acquire, in a current traffic environment, current image traffic data collected by a preset image capture device during a current period and historical text traffic data stored during historical periods, wherein the image capture device is a plurality of drone devices deployed in the current traffic environment to perform joint observation of the current traffic environment; The detection module 102 is used to perform data analysis and processing on the current image traffic data based on a preset multi-level modeling network to obtain the current traffic flow characteristics and generate the traffic flow detection results of the current traffic environment in the current time period based on the current traffic flow characteristics; The prediction module 103 is used to perform data analysis and processing on the current image traffic data and the historical text traffic data based on a preset multimodal encoding network to obtain predicted traffic flow characteristics, and generate traffic flow prediction results for the current traffic environment in the future time period based on the predicted traffic flow characteristics.

[0124] In an exemplary embodiment, the detection module 102 is also used to: based on the first type of modeling layer, perform data feature extraction processing on the current image traffic data to obtain feature data of different scales output by the backbone structure; based on the second type of modeling layer, perform fusion processing on the feature data of different scales to obtain the output of the neck structure, and integrate the output of the neck structure into the current traffic flow feature.

[0125] In an exemplary embodiment, the detection module 102 is also used to: based on the spatial relationship modeling layer, perform spatial relationship modeling on different pixels in the current image traffic data to obtain first feature data that structuredly reflects the spatial relationship; based on the channel relationship modeling layer, perform channel relationship modeling on different channels in the first feature data to obtain second feature data that structuredly reflects the channel relationship; use the first feature data and the second feature data as feature data of different scales outputted by the backbone structure; based on the semantic relationship modeling layer, perform semantic relationship modeling on different semantic objects in the first feature data and the second feature data to obtain the output of the neck structure that structuredly reflects the semantic relationship, and integrate the output of the neck structure into the current traffic flow feature.

[0126] In an exemplary embodiment, the detection module 102 is also used to: perform semantic relationship modeling on the second feature data based on the first semantic relationship modeling layer to obtain first intermediate feature data; perform semantic relationship modeling on the first feature data and the first intermediate feature data based on the second semantic relationship modeling layer to obtain second intermediate feature data; perform semantic relationship modeling on the second feature data and the second intermediate feature data based on the third semantic relationship modeling layer to obtain third intermediate feature data; and use the second intermediate feature data and the third intermediate feature data as outputs of the neck structure.

[0127] In an exemplary embodiment, the prediction module 103 is also used to: based on the text coding layer, perform text feature extraction processing on historical text traffic data to obtain historical traffic flow text features; based on the visual coding layer, perform visual feature extraction processing on current image traffic data to obtain current traffic flow visual features; based on the text visual coding layer, perform data alignment processing on current traffic flow visual features and historical traffic flow text features at the text data level, visual data level and text visual data level respectively to obtain aligned data corresponding to each data level; perform fusion processing on the aligned data corresponding to each data level to obtain predicted traffic flow features.

[0128] In an exemplary embodiment, the prediction module 103 is also used to: perform text feature extraction processing on historical text traffic data based on the text encoding layer to obtain initial historical traffic flow text features; perform forward time series modeling processing and reverse time series modeling processing on the initial historical traffic flow text features based on the bidirectional long and short time series modeling layer to obtain historical traffic flow text features.

[0129] In an exemplary embodiment, the prediction module 103 is also used to: based on the historical traffic flow text features, perform data alignment processing on the current traffic flow visual features and the historical traffic flow text features at the text data level through the text visual coding layer to obtain first aligned data at the text data level; based on the current traffic flow visual features, perform data alignment processing on the current traffic flow visual features and the historical traffic flow text features at the visual data level through the text visual coding layer to obtain second aligned data at the visual data level; based on the splicing features corresponding to the historical traffic flow text features and the current traffic flow visual features, perform data alignment processing on the splicing features at the text visual data level through the text visual coding layer to obtain third aligned data at the text visual data level; and use the first aligned data, the second aligned data, and the third aligned data as the aligned data corresponding to each data level.

[0130] Each module in the aforementioned traffic flow detection and prediction device based on drone feedback collaborative optimization can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0131] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in any of the above embodiments when executing the computer program.

[0132] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any of the above embodiments are implemented.

[0133] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0134] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0135] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A traffic flow detection and prediction method based on UAV feedback collaborative optimization, characterized in that: The method comprises: In a current traffic environment, current image traffic data collected by a preset image capturing device in a current time period and historical text traffic data stored in a historical time period are obtained, wherein the image capturing device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to perform joint observation of the current traffic environment; Based on a preset multi-level modeling network, the current image traffic data is subjected to data analysis processing to obtain current traffic flow characteristics, and a traffic flow detection result of the current traffic environment in the current time period is generated according to the current traffic flow characteristics; Based on a preset multimodal encoding network, data analysis and processing are performed on the current image traffic data and the historical text traffic data to obtain predicted traffic flow characteristics, and traffic flow prediction results for the current traffic environment in future time periods are generated based on the predicted traffic flow characteristics.

2. The method according to claim 1, characterized in that The multi-level modeling network includes a first type of modeling layer in the backbone structure and a second type of modeling layer in the neck structure; The current image traffic data is subjected to data analysis and processing based on the preset multi-level modeling network to obtain current traffic flow characteristics, including: Based on the first type of modeling layer, data feature extraction processing is performed on the current image traffic data to obtain feature data of different scales output by the backbone structure; Based on the second type of modeling layer, feature data of different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow feature.

3. The method according to claim 2, characterized in that The first type of modeling layer includes a spatial relationship modeling layer and a channel relationship modeling layer, and the second type of modeling layer includes a semantic relationship modeling layer; The data feature extraction process is performed on the current image traffic data based on the first type of modeling layer to obtain feature data of different scales output by the backbone structure, including: Based on the spatial relationship modeling layer, spatial relationship modeling processing is performed on different pixels in the current image traffic data to obtain first feature data that structurally reflects the spatial relationship; Based on the channel relationship modeling layer, performing channel relationship modeling processing on different channels in the first feature data to obtain second feature data that structurally reflects the channel relationship; Using the first feature data and the second feature data as feature data of different scales output by the backbone structure; Based on the second modeling layer, the feature data of different scales are fused to obtain the output of the neck structure, and the output of the neck structure is integrated into the current traffic flow feature, including: Based on the semantic relationship modeling layer, semantic relationship modeling is performed on different semantic objects in the first feature data and the second feature data to obtain a structured output of a neck structure that reflects the semantic relationship, and the output of the neck structure is integrated into the current traffic flow feature.

4. The method according to claim 3, characterized in that The semantic relationship modeling layer includes a first semantic relationship modeling layer, a second semantic relationship modeling layer and a third semantic relationship modeling layer; The step of performing semantic relationship modeling on different semantic objects in the first feature data and the second feature data based on the semantic relationship modeling layer to obtain a structured output of the neck structure reflecting the semantic relationship includes: Based on the first semantic relationship modeling layer, performing semantic relationship modeling processing on the second feature data to obtain first intermediate feature data; Based on the second semantic relationship modeling layer, performing semantic relationship modeling processing on the first feature data and the first intermediate feature data to obtain second intermediate feature data; Based on the third semantic relationship modeling layer, performing semantic relationship modeling processing on the second feature data and the second intermediate feature data to obtain third intermediate feature data; The second intermediate feature data and the third intermediate feature data are used as outputs of the neck structure.

5. The method according to claim 1, wherein The multimodal encoding network includes a text encoding layer, a visual encoding layer and a text-visual encoding layer; The preset multimodal encoding network is used to perform data analysis on the current image traffic data and the historical text traffic data to obtain predicted traffic flow characteristics, including: Based on the text encoding layer, performing text feature extraction processing on the historical text traffic data to obtain historical traffic flow text features; Based on the visual coding layer, performing visual feature extraction processing on the current image traffic data to obtain current traffic flow visual features; Based on the text visual coding layer, data alignment processing is performed on the current traffic flow visual features and the historical traffic flow text features at the text data level, the visual data level, and the text visual data level to obtain aligned data corresponding to each data level; The aligned data corresponding to each data level are fused to obtain the predicted traffic flow characteristics.

6. The method according to claim 5, characterized in that The multimodal encoding network also includes a bidirectional long-short time series modeling layer; The text feature extraction process is performed on the historical text traffic data based on the text encoding layer to obtain historical traffic flow text features, including: Based on the text encoding layer, performing text feature extraction processing on the historical text traffic data to obtain initial historical traffic flow text features; Based on the bidirectional long-short time series modeling layer, forward time series modeling processing and reverse time series modeling processing are performed on the initial historical traffic flow text features to obtain historical traffic flow text features.

7. The method according to claim 5, characterized in that Based on the text visual coding layer, the current traffic flow visual features and the historical traffic flow text features are aligned at the text data level, the visual data level, and the text visual data level to obtain aligned data corresponding to each data level, including: Based on the historical traffic flow text features, the current traffic flow visual features and the historical traffic flow text features are aligned at the text data level through the text visual coding layer to obtain first aligned data at the text data level; Based on the current traffic flow visual feature as a reference, the current traffic flow visual feature and the historical traffic flow text feature are aligned at the visual data level through the text visual coding layer to obtain second aligned data at the visual data level; Based on the splicing features corresponding to the historical traffic flow text features and the current traffic flow visual features, data alignment processing is performed on the splicing features at the text visual data level through the text visual coding layer to obtain third alignment data at the text visual data level; The first aligned data, the second aligned data, and the third aligned data are used as aligned data corresponding to each data level.

8. A traffic flow detection and prediction device based on UAV feedback collaborative optimization, characterized in that: The device comprises: An acquisition module is configured to acquire, in a current traffic environment, current image traffic data collected by a preset image capture device during a current period and historical text traffic data stored during historical periods, wherein the image capture device is a plurality of unmanned aerial vehicle devices deployed in the current traffic environment to perform joint observation of the current traffic environment; a detection module configured to perform data analysis on the current image traffic data based on a preset multi-level modeling network to obtain current traffic flow characteristics, and generate a traffic flow detection result of the current traffic environment in the current time period based on the current traffic flow characteristics; The prediction module is used to perform data analysis and processing on the current image traffic data and the historical text traffic data based on a preset multimodal encoding network to obtain predicted traffic flow characteristics, and generate traffic flow prediction results for the current traffic environment in the future time period based on the predicted traffic flow characteristics.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Internet unmanned aerial vehicle capable of achieving continuous endurance

    CN105652886A

  • Data processing method and device, electronic equipment and storage medium

    CN115115913A

  • Pedestrian flow prediction system and method

    CN117592599A

  • Cross-terminal finger vein recognition technology based on graph neural network

    CN119229486A

  • Mama-based traffic flow prediction method

    CN119942813A