A vehicle cross-scene matching method and system based on front and rear vehicle dynamic topology
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AI SUPER EYE TECH CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]本申请提供一种基于前后车动态拓扑的车辆跨场景匹配方法及系统,用于针对解决现有技术中车辆外观识别受光照、视角和遮挡影响,轨迹关联在密集交通流中容易失效的技术问题
Smart Images

Figure CN122530628A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation technology, specifically to a method and system for cross-scenario vehicle matching based on the dynamic topology of preceding and following vehicles. Background Technology
[0002] Vehicle appearance recognition technology typically relies on matching features such as vehicle color, texture, and model. However, these features are easily affected by different lighting conditions, different viewing angles, and when the vehicle is partially occluded, leading to a decrease in recognition accuracy. Furthermore, trajectory-based association methods rely on matching vehicle motion trajectories, but in dense traffic and complex vehicle behavior, vehicle trajectories are prone to intersections or interruptions, resulting in matching failures. Summary of the Invention
[0003] This application provides a vehicle cross-scene matching method and system based on the dynamic topology of front and rear vehicles, which is used to address the technical problems in the prior art where vehicle appearance recognition is affected by lighting, viewing angle and occlusion, and trajectory association is prone to failure in dense traffic flow.
[0004] In view of the above problems, this application provides a vehicle cross-scene matching method and system based on the dynamic topology of front and rear vehicles.
[0005] The first aspect of this application provides a vehicle cross-scene matching method based on dynamic topology of front and rear vehicles, the method comprising:
[0006] A real-time image stream is acquired using pre-deployed image acquisition devices, comprising a first image stream acquired by a first device and a second image stream acquired by a second device. A first image corresponding to the target vehicle is filtered from the first image stream, and the first image is analyzed to obtain first target information of the target vehicle. A target image is determined based on the second image stream, and the target image is analyzed to obtain second target information. The first target information and the second target information are fused and matched based on a cross-scene multimodal fusion strategy to obtain a preliminary matching result. If the preliminary matching result meets predetermined constraints, an alignment command is issued. Based on the alignment command, cross-viewpoint alignment processing is performed on the first device and the second device to obtain global target information of the target vehicle.
[0007] A second aspect of this application provides a vehicle cross-scenario matching system based on dynamic topology of front and rear vehicles, the system comprising:
[0008] An image acquisition module is used to acquire a real-time image stream through a pre-deployed image acquisition device, wherein the real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device; a first analysis module is used to filter a first image corresponding to a target vehicle in the first image stream and analyze the first image to obtain first target information of the target vehicle; a second analysis module is used to determine a target image based on the second image stream and analyze the target image to obtain second target information; a fusion matching module is used to fuse and match the first target information and the second target information based on a cross-scene multimodal fusion strategy to obtain a preliminary matching result; an instruction sending module is used to issue an alignment instruction if the preliminary matching result meets predetermined constraints; and a processing module is used to perform cross-view alignment processing on the first device and the second device based on the alignment instruction to obtain global target information of the target vehicle.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] This application acquires a real-time image stream using pre-deployed image acquisition devices. The real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device. A first image corresponding to a target vehicle is selected from the first image stream, and the first image is analyzed to obtain first target information of the target vehicle. A target image is determined based on the second image stream, and the target image is analyzed to obtain second target information. The first target information and the second target information are fused and matched based on a cross-scene multimodal fusion strategy to obtain a preliminary matching result. If the preliminary matching result meets predetermined constraints, an alignment command is issued. Based on the alignment command, the first device and the second device undergo cross-viewpoint alignment processing to obtain global target information of the target vehicle. This invention solves the technical problems in the prior art where vehicle appearance recognition is affected by illumination, viewing angle, and occlusion, and trajectory association easily fails in dense traffic flow. By utilizing the dynamic topology between vehicles for feature fusion, it achieves the technical effect of improving the accuracy and robustness of cross-camera and cross-scene matching in complex traffic environments. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1A schematic diagram of a vehicle cross-scene matching method based on the dynamic topology of front and rear vehicles provided in this application embodiment;
[0013] Figure 2 This is a schematic diagram of a vehicle cross-scenario matching system based on the dynamic topology of the front and rear vehicles, provided as an embodiment of this application.
[0014] Explanation of reference numerals in the attached figures: Image acquisition module 11, First analysis module 12, Second analysis module 13, Fusion matching module 14, Instruction sending module 15, Processing module 16. Detailed Implementation
[0015] This application provides a vehicle cross-scene matching method and system based on the dynamic topology of front and rear vehicles. It addresses the technical problems in the prior art where vehicle appearance recognition is affected by lighting, viewing angle and occlusion, and trajectory association is prone to failure in dense traffic flow. By utilizing the dynamic topology structure between vehicles for feature fusion, it achieves the technical effect of improving the accuracy and robustness of cross-camera and cross-scene matching in complex traffic environments.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0017] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0018] Example 1, as Figure 1 As shown, this application provides a vehicle cross-scene matching method based on the dynamic topology of front and rear vehicles, the method comprising:
[0019] Step S100: Acquire a real-time image stream through a pre-deployed image acquisition device, wherein the real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device.
[0020] In this embodiment, when acquiring a real-time image stream using a pre-deployed image acquisition device, the image acquisition device is first deployed according to a deployment strategy. Then, the device is activated to acquire data according to the acquisition strategy to obtain the real-time image stream. The real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device.
[0021] Furthermore, the method provided in the application embodiments, which acquires a real-time image stream through a pre-deployed image acquisition device, further includes:
[0022] Read the deployment strategy and deploy the image acquisition device according to the deployment strategy; read the acquisition strategy and activate the image acquisition device to acquire data according to the acquisition strategy to obtain the real-time image stream.
[0023] Furthermore, the method provided in the application embodiments also includes:
[0024] The deployment strategy refers to deploying the image acquisition equipment at intervals of 50 meters to 200 meters along the monitored road section on the side of the road, with an installation height of 3 meters to 6 meters and a field of view covering 3 to 6 lanes. The image acquisition equipment refers to a monocular camera or a binocular camera.
[0025] Furthermore, the method provided in the application embodiments also includes:
[0026] The acquisition strategy refers to achieving data synchronization between cameras within the array through the PTP time synchronization protocol, with an acquisition frame rate of 25-30fps and a synchronization error of ≤10ms.
[0027] In this embodiment, to acquire a real-time image stream, the image acquisition devices are first deployed according to a deployment strategy. The deployment strategy specifies that image acquisition devices are deployed at intervals of 50 to 200 meters along the monitored road section on the roadside, with an installation height of 3 to 6 meters to ensure that the field of view covers 3 to 6 lanes. This deployment scheme achieves real-time monitoring of different lanes by using monocular or binocular cameras.
[0028] Next, the acquisition strategy is read and the image acquisition device is activated to acquire data according to the strategy. The acquisition strategy includes synchronizing camera data within the array through the Precision Time Protocol (PTP). The PTP protocol ensures precise synchronization of data acquisition time between devices, reducing data deviations caused by time differences. The acquisition frame rate is set to 25-30fps, and the synchronization error does not exceed 10 milliseconds to ensure the continuity and real-time performance of the image stream. These real-time image streams include a first image stream acquired by the first device and a second image stream acquired by the second device, forming a multi-view image stream.
[0029] Through this high-precision time synchronization, the image acquisition device can stably output a real-time image stream. Specifically, the real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device. Each device acquires images at a specified frame rate of 25-30 fps, with synchronization errors controlled within 10 milliseconds. This data acquisition setup ensures a high-frequency, low-latency image data stream, meeting the needs of subsequent image analysis, vehicle perception, and target matching.
[0030] Through the above steps, the image acquisition device acquires real-time image streams at different locations.
[0031] Step S200: Filter the first image corresponding to the target vehicle in the first image stream, and analyze the first image to obtain the first target information of the target vehicle.
[0032] In this embodiment, the first image corresponding to the target vehicle is first filtered from the first image stream. During this process, the YOLO series of object detection algorithms are used for vehicle detection. The YOLO algorithm extracts features from the image using a convolutional neural network (CNN), identifies vehicles in the image, and generates bounding boxes for each vehicle, marking its specific location. Then, for each frame in the first image stream, it is checked whether a vehicle appears. If one or more vehicles are present in the image, the image frame is marked as containing vehicles, and this image is used as the first image.
[0033] Next, we analyze the first image. In this process, we first decode the first image using depth estimation techniques to obtain the first spatiotemporal state sequence of the target vehicle. This spatiotemporal state sequence includes the target vehicle's three-dimensional coordinates, three-dimensional dimensions, velocity vector, and heading angle. Then, this spatiotemporal state information is added to the target vehicle's first target information.
[0034] Furthermore, in the method provided in the application embodiments, filtering the first image corresponding to the target vehicle in the first image stream and analyzing the first image to obtain the first target information of the target vehicle further includes:
[0035] The first image is solved based on depth estimation technology to obtain a first spatiotemporal state sequence; the first spatiotemporal state sequence is added to the first target information; wherein, the first spatiotemporal state sequence includes a first three-dimensional coordinate, a first three-dimensional size, a first velocity vector and a first heading angle.
[0036] In this embodiment, the first image is first processed using depth estimation technology to obtain the first spatiotemporal state sequence of the target vehicle. Specifically, the Zhang Zhengyou calibration method is used to calibrate the binocular camera deployed on the roadside, obtaining the camera's internal parameters (including focal length f and principal point coordinates) and distortion coefficients. , , , , The parameters include the rotation matrix R and translation vector T between the binocular systems. Based on these parameters, distortion correction is performed on the left and right views to eliminate radial and tangential lens distortion. The correction formula is as follows: , .
[0037] in, = + (x, y) are the corrected ideal image coordinates. The coordinates are those of the distorted image. Then, the Bouguet epipolar correction algorithm is used to reproject the left and right images onto the same plane, ensuring that corresponding points lie on the same horizontal line, thus providing geometrically constrained input for stereo matching.
[0038] On the corrected left and right images, a stereo matching method based on deep neural networks, such as PSMNet, is applied for pixel-by-pixel matching to calculate the difference in the horizontal coordinates of corresponding points in the left and right views, generating a dense disparity map, where the disparity value of each pixel is denoted as disparity. Based on the triangulation principle of binocular vision, and combined with the calibrated camera baseline length b and focal length f, the disparity values are converted into depth values using a depth calculation formula. This process generates a depth map, enabling depth perception of the scene. This step achieves a quantitative mapping from a two-dimensional image to three-dimensional spatial depth. Where d is the distance from the vehicle to the camera, b is the baseline of the stereo camera, f is the focal length of the stereo camera, and disparity is the disparity value of the stereo image.
[0039] Based on the depth map and pixel coordinates, the pixel points (u, v) and their corresponding disparities d in the image coordinate system are transformed to the world coordinate system using the reprojection matrix Q. The reprojection matrix Q is derived from the camera calibration parameters and has the following form: .in( () represents the coordinates of the principal point of the left camera. Let x be the x-coordinate of the principal point of the right camera. Let f be the ordinate of the principal point of the right camera, f be the focal length, and b be the baseline length. The conversion formula is as follows: .
[0040] The final three-dimensional coordinates are (X / W, Y / W, Z / W), from which the first three-dimensional coordinates (X, Y, Z) of the target vehicle are calculated. At the same time, combined with the two-dimensional bounding box and depth information output by the target detection network, the first three-dimensional dimensions (length, width, height) of the vehicle are calculated, forming an accurate description of the vehicle's outline.
[0041] By repeating the above calculation process on multiple consecutive frames of images, the first three-dimensional coordinates of the target vehicle in each frame are obtained. The displacement vector is calculated using inter-frame difference and timestamps, and then divided by the time interval to obtain the first velocity vector. This vector describes the vehicle's speed and direction of motion in three-dimensional space. Simultaneously, based on the main direction or vehicle orientation change fitted from the three-dimensional bounding box in consecutive frames, a first heading angle is determined, which defines the vehicle's driving posture in the global coordinate system.
[0042] The first three-dimensional coordinates, first three-dimensional dimensions, first velocity vector, and first heading angle obtained from the frame-by-frame calculation are organized and stored in chronological order to form a first spatiotemporal state sequence that can completely describe the motion state of the target vehicle in the continuous time domain. This sequence is then added to the first target information as a structured data unit.
[0043] Furthermore, in the method provided in the application embodiments, after adding the first spatiotemporal state sequence to the first target information, it further includes:
[0044] Based on the first image, a first target perception domain is determined with the target vehicle as the center, and a first star-shaped topology map of the first target perception domain is drawn; the first star-shaped topology map is encoded to obtain a first topological feature vector, and added to the first target information; wherein, any edge in the first star-shaped topology map has at least the identifiers of a first relative distance, a first relative speed, a first relative angle, a first TTC (time of collision), and a first lane relationship.
[0045] In this embodiment, a first target perception domain is first determined based on the first image, centered on the target vehicle. This perception domain is defined according to a dynamic distance threshold and lane constraints. A rectangular area extending 50 meters forward and 30 meters backward from the target vehicle, and laterally covering adjacent left and right lanes, is used as the baseline perception range. Simultaneously, the scaling factor of this range is dynamically adjusted according to actual traffic scenarios (such as highways or congested sections). This is achieved through spatial coordinate calculation and lane-level mapping. Specifically, using the calculated first three-dimensional coordinates and high-precision map lane line data, it is determined whether other vehicles fall within this dynamic rectangular area, thereby obtaining a local spatial range containing the target vehicle and all its neighboring vehicles, i.e., the first target perception domain.
[0046] Within the defined first target perception domain, a first star-shaped topology graph is drawn, with the target vehicle as the central node and all neighboring vehicles as leaf nodes. This topology graph is a star structure with the target vehicle as the root node and neighboring vehicles as child nodes. Each edge in the graph connects the target vehicle to a neighboring vehicle. During the drawing process, each edge is assigned multi-dimensional relationship features as identifiers, specifically including: a first relative distance, i.e., the Euclidean distance between the target vehicle and neighboring vehicles in the world coordinate system; a first relative velocity, i.e., the magnitude of the difference between their velocity vectors; a first relative angle, i.e., the absolute value of the difference between their heading angles; a first TTC (Time to Collision), calculated by dividing the first relative distance by the first relative velocity; and a first lane relationship, which identifies whether the target vehicle and neighboring vehicles are in the same lane, the left lane, or the right lane by comparing their lateral positions and lane line information. Thus, a graph structure data centered on the target vehicle with edges bearing multi-dimensional relationship features is obtained, i.e., the first star-shaped topology graph.
[0047] After drawing the first star-shaped topology, the inductive graph neural network GraphSAGE is used to encode the topology to generate the first topological feature vector. During the encoding process, the center node of the target vehicle is first used as the anchor point, and all its neighboring nodes are sampled. Then, the feature vectors of the neighboring nodes are aggregated using the mean aggregation function of GraphSAGE. Simultaneously, the multidimensional features of each edge (such as the first relative distance and the first relative velocity) are embedded as edge features in the aggregation process to enhance the expression of structural information. After two layers of aggregation, the center node finally outputs a fixed-dimensional embedding vector, i.e., the first topological feature vector. This vector encodes the dynamic topological structure information formed by the target vehicle and its surrounding vehicles, and through symmetric aggregation functions and relative feature design in the local coordinate system, invariance to node arrangement order and viewpoint rotation is achieved. Finally, the generated first topological feature vector is added as a structured feature unit to the first target information. Thus, the first target information, based on the original first spatiotemporal state sequence, adds topological features describing the interaction relationships between vehicles, forming a complete multimodal description of the target vehicle.
[0048] Step S300: Determine the target image based on the second image stream, and analyze the target image to obtain the second target information.
[0049] In this embodiment, when determining the target image based on the second image stream, the spatial distance between the first and second devices is first calculated. Then, the second image stream is filtered by combining the target vehicle's average speed and the calculated spatial distance to obtain a candidate image set. Next, any candidate image is extracted from the candidate image set, and the target information in that image is obtained. Finally, by comparing the similarity between the target information and the first target information, the image corresponding to the target information with the highest similarity is selected as the target image.
[0050] Next, the target image is analyzed. First, the same depth estimation technique as described above is used to solve the target image. Based on the binocular stereo vision method, after distortion correction and epipolar correction of the left and right views, a dense disparity map is generated using a stereo matching algorithm based on deep neural networks (such as PSMNet). According to the triangulation principle, the disparity values are converted into depth values, and the pixel coordinates are converted into three-dimensional coordinates in the world coordinate system through the reprojection matrix Q, thereby obtaining the second three-dimensional coordinates and second three-dimensional dimensions of the target vehicle from the perspective of the second device. This process is repeated for multiple consecutive frames and combined with inter-frame difference to calculate the second velocity vector and the second heading angle. These are organized in chronological order to form a second spatiotemporal state sequence. Second, with the target vehicle as the center, a second target perception domain is defined based on the dynamic distance threshold and lane constraints. Neighboring vehicles within the perception domain are used as nodes to construct a second star-shaped topology graph. Each edge in the graph is assigned multi-dimensional features such as the second relative distance, the second relative velocity, the second relative angle, the second TTC, and the second lane relationship. The topological graph is encoded using a graph neural network, GraphSAGE. Neighbor features are aggregated by mean aggregation and edge features are incorporated to generate a second topological feature vector with topological invariance. Finally, the second spatiotemporal state sequence and the second topological feature vector are integrated together to form the second target information.
[0051] Furthermore, in the method provided in the application embodiments, determining the target image based on the second image stream further includes:
[0052] Calculate the spatial distance between the first device and the second device; combine the target average speed of the target vehicle and the spatial distance to filter the second image stream and obtain a candidate image set; extract any candidate image from the candidate image set and obtain any target information of the arbitrary candidate image; determine the target image with the highest similarity between the arbitrary target information and any information of the first target information.
[0053] In this embodiment, the spatial distance between the first device and the second device is first calculated. When image acquisition devices are deployed on the roadside, the world coordinates of each device are measured or read from a high-precision map, and the Euclidean distance formula is used to calculate the Euclidean distance between the first device and the second device in three-dimensional space to obtain the spatial distance.
[0054] After obtaining the spatial distance, the motion characteristics of the target vehicle exhibited in the first image stream, i.e., the target average speed, are combined. This average speed is calculated by the average of the first velocity vector in the first spatiotemporal state sequence on the time axis, reflecting the typical driving speed of the target vehicle on the road segment. The spatial distance is divided by the target average speed to estimate the estimated travel time required for the target vehicle to travel from the first device's field of view to the second device's field of view. Based on this time window, the timestamps of the second image stream are retrieved, and all image frames located within this time window are selected. These image frames together constitute the candidate image set.
[0055] For each frame in the candidate image set, the same perception and analysis methods as described above are executed to obtain arbitrary target information of all vehicles in each frame. Specifically, for each arbitrary candidate image frame, depth estimation technology is first used for calculation. Disparity maps are generated through distortion correction, epipolar correction, and stereo matching, and then converted into depth maps. The reprojection matrix Q is used to calculate the three-dimensional coordinates and three-dimensional dimensions of each vehicle in the image. Differential operations are performed on consecutive frames to obtain velocity vectors and heading angles, forming the spatiotemporal state sequence of the vehicle. Simultaneously, a perception domain is defined with the vehicle as the center, and a star-shaped topology map is constructed. Each edge in the map is assigned features such as relative distance, relative speed, relative angle, TTC, and lane relationship. The graph neural network GraphSAGE is used to encode the topology map to generate the topological feature vector of the vehicle. The spatiotemporal state sequence and the topological feature vector are integrated to form arbitrary target information describing the complete motion state and interaction relationships of the vehicle from the current perspective.
[0056] Finally, the unique target image is determined from the candidate image set based on the criterion of the highest similarity between any target information and the first target information. The similarity calculation employs a multimodal fusion strategy, comprehensively measuring the consistency of the two in terms of topological structure and motion state. First, the cosine similarity between the topological feature vector in any target information and the first topological feature vector in the first target information is calculated, denoted as . This value quantifies the similarity of the target vehicle's local dynamic topology between the second and first viewpoints. Secondly, the consistency of motion states is quantified by extracting velocity vectors from the spatiotemporal state sequences of any target information. With heading angle Extract the velocity vector from the first spatiotemporal state sequence in the first target information. With heading angle Calculate the cosine similarity of velocity vectors. Simultaneously calculate the normalized score of the absolute value of the heading angle difference. The two are weighted and fused to obtain the motion state consistency score. ,in The weights are preset. Finally, the overall similarity score is obtained by weighted summation. ,in To adjust the fusion coefficients contributing to topological features and motion state, the image frame corresponding to the arbitrary target information of all vehicles in the candidate image set with the highest comprehensive similarity score to the first target information is selected as the final target image.
[0057] Furthermore, in the method provided in the application embodiments, before fusing and matching the first target information and the second target information based on a cross-scenario multimodal fusion strategy to obtain a preliminary matching result, the method further includes:
[0058] The scene information is analyzed according to the cross-scene multimodal fusion strategy to obtain the matching confidence; the first target information and the second target information are fused and matched based on the matching confidence.
[0059] In this embodiment, before fusing and matching the first target information and the second target information based on the cross-scene multimodal fusion strategy, the scene information is first analyzed according to the cross-scene multimodal fusion strategy to obtain the matching confidence. Specifically, the current frame image is extracted from the real-time image stream acquired by the first device or the second device, scaled to 224×224 pixels, and then input into a pre-trained convolutional neural network for scene encoding. The network uses a ResNet-18 model pre-trained on the ImageNet dataset. The feature map output from the last convolutional layer is taken and subjected to global average pooling to obtain a 512-dimensional feature vector. Then, it is compressed to 128 dimensions through a fully connected layer as a scene confidence vector representing contextual information such as the current scene's illumination intensity, weather conditions, and occlusion degree. The training of the network adopts a self-supervised contrastive learning approach: in a large number of cross-camera image pairs, images of the same scene at different times have similar confidence encodings, while the encodings of different scenes are far apart, thereby learning feature representations that can reflect the difficulty of scene matching. Thus, a 128-dimensional scene confidence vector was obtained, which is used to quantify the quality and reliability of the current matching environment.
[0060] After obtaining the scene confidence level, the first target information and the second target information are fused and matched based on the confidence level.
[0061] Step S400: Based on the cross-scenario multimodal fusion strategy, the first target information and the second target information are fused and matched to obtain a preliminary matching result.
[0062] In this embodiment, when fusing and matching the first target information and the second target information based on a cross-scene multimodal fusion strategy, the aforementioned scene confidence score is first extracted, and the fusion and matching of the first target information and the second target information is performed based on this confidence score. First, a first topological feature vector is extracted from the first target information, and a second topological feature vector is extracted from the second target information. The cosine similarity between the two is calculated to obtain the topological feature similarity. Simultaneously, a first velocity vector and a first heading angle are extracted from the first spatiotemporal state sequence in the first target information, and a second velocity vector and a second heading angle are extracted from the second spatiotemporal state sequence in the second target information. The cosine similarity of the velocity vectors is then calculated. And calculate the normalized score of the absolute value of the heading angle difference. The motion state consistency score is obtained by weighted summation of the two. .
[0063] Subsequently, the topological feature similarity was... Consistency score of motion state The aforementioned 128-dimensional scene confidence vector is concatenated into an input vector, which is then fed into a lightweight adaptive fusion network for processing. This network employs a three-layer multilayer perceptron structure, with an input layer dimension of 1+1+128=130, hidden layer dimensions of 64 and 32 respectively, ReLU activation function, and a single output neuron that outputs a comprehensive matching confidence score between 0 and 1 using the sigmoid function. The network learns through supervised training to dynamically adjust the contribution weights of topological features and motion state features. The training dataset consists of a large number of cross-camera matched and unmatched pairs, with each sample containing... , The network uses a scene confidence vector with a label of 1 for matching and 0 for non-matching. The loss function is binary cross-entropy loss, and the optimizer is Adam. During training, the multilayer perceptron parameters are updated through backpropagation, enabling the network to adaptively fuse the two similarities based on scene confidence and output a more accurate matching score.
[0064] Finally, the overall matching confidence score will be used. Compare with a preset matching threshold T (e.g., 0.7). If If the value is greater than or equal to T, then the first target information and the second target information are determined to correspond to the same target vehicle, thus forming a preliminary matching result.
[0065] Step S500: If the preliminary matching result meets the predetermined constraints, then an alignment command is issued.
[0066] In this embodiment of the application, if the preliminary matching result meets the predetermined constraints, that is, the first target information and the second target information correspond to the same target vehicle, an alignment command is issued to trigger subsequent cross-view alignment processing.
[0067] Step S600: Perform cross-view alignment processing on the first device and the second device based on the alignment instruction to obtain the target global information of the target vehicle.
[0068] In this embodiment, when performing cross-view alignment processing on the first and second devices based on alignment instructions, the first vehicle trajectory of the target vehicle collected by the first device is first acquired, and the second vehicle trajectory of the target vehicle collected by the second device is also acquired. Then, based on a predetermined algorithm and constrained by the logical consistency of macroscopic traffic flow, a graph matching algorithm or a consistency propagation algorithm is used to obtain the target vehicle's continuous spatiotemporal trajectory. Finally, the obtained continuous spatiotemporal trajectory is used as the target's global information.
[0069] Furthermore, in the method provided in the application embodiment, the cross-view alignment processing of the first device and the second device based on the alignment instruction to obtain the target global information of the target vehicle further includes:
[0070] The first vehicle trajectory of the target vehicle collected by the first device is obtained; the second vehicle trajectory of the target vehicle collected by the second device is obtained; based on a predetermined algorithm and constrained by the logical consistency of macroscopic traffic flow, the target continuous spatiotemporal trajectory of the target vehicle is obtained, wherein the predetermined algorithm refers to a graph matching algorithm or a consistency propagation algorithm; the target continuous spatiotemporal trajectory is used as the target global information.
[0071] In this embodiment, firstly, a continuous position sequence of the target vehicle within the field of view of a single camera is extracted from the image stream of the first device to obtain a first vehicle trajectory. This trajectory is formed by connecting the calculated first three-dimensional coordinates in chronological order, describing the movement path of the target vehicle within the road segment monitored by the first device. Simultaneously, a continuous position sequence of the target vehicle within the field of view of another camera is extracted from the image stream of the second device to obtain a second vehicle trajectory. This trajectory is formed by connecting the calculated second three-dimensional coordinates in chronological order, describing the movement path of the target vehicle within the road segment monitored by the second device.
[0072] Next, based on a predetermined algorithm, the trajectories of the first and second vehicles are globally correlated and stitched together. Constrained by the logical consistency of macroscopic traffic flow, the target vehicle's continuous spatiotemporal trajectory is obtained. The predetermined algorithm can be either a graph matching algorithm or a consistency propagation algorithm. Both algorithms achieve global optimization matching across camera trajectories in different ways, ensuring that the matching results conform to the physical laws of real traffic scenarios.
[0073] If a graph matching algorithm is used, the Hungarian algorithm will be used as an example. To achieve globally optimal matching, the trajectory sets of all vehicles within the field of view of the first device and the field of view of the second device are obtained simultaneously, with the target vehicle's trajectory included in both sets. First, a matching cost matrix is constructed, with the rows of the matrix representing each vehicle trajectory in the first device and the columns representing each vehicle trajectory in the second device. The element in the i-th row and j-th column of the matrix represents the matching cost between the i-th trajectory in the first device and the j-th trajectory in the second device. The matching cost is calculated as follows: First, a first topological feature vector is extracted from the target information corresponding to the i-th trajectory; second, a second topological feature vector is extracted from the target information corresponding to the j-th trajectory. The cosine similarity between the two vectors is calculated, with a value closer to 1 indicating greater topological similarity. Simultaneously, velocity vectors and heading angles are extracted from the spatiotemporal state sequences corresponding to the two trajectories. The cosine similarity between the two velocity vectors is calculated, and the absolute value of the difference between the two heading angles is divided by pi. The normalized score of this ratio is then subtracted from 1. The velocity similarity and heading angle similarity are weighted and summed to obtain the motion state consistency score. Finally, the topological feature cosine similarity and motion state consistency score are weighted and summed to obtain the comprehensive similarity. This comprehensive similarity is then subtracted from 1 to obtain the final matching cost; a higher comprehensive similarity results in a lower matching cost. This process is repeated for all rows and columns to fill the entire cost matrix. After constructing the cost matrix, the Hungarian algorithm is run to solve it. Specifically, the cost matrix is first reduced by rows, meaning the minimum value of each row is subtracted from all elements in that row, ensuring each row contains at least one zero element. Then, the cost matrix is reduced by columns, meaning the minimum value of each column is subtracted from all elements in that column, ensuring each column also contains at least one zero element. Next, all zero elements in the matrix are covered with the minimum number of horizontal or vertical lines. If the required number of lines is less than the number of rows in the matrix, the minimum value among the uncovered elements is found, subtracted from the uncovered elements, and added to the elements at the line intersections. Zero-element coverage is then repeated until the required number of lines equals the number of rows in the matrix. At this point, the initial optimal matching scheme can be determined from the position of the zero elements. After obtaining candidate matching schemes in each iteration, a logical consistency constraint based on macroscopic traffic flow is introduced for verification. This constraint requires that the relative order of vehicles in the fields of view of different cameras must remain consistent. For any two trajectories in the first device, if one trajectory is determined to precede the other based on its entry time or spatial position into the field of view, then the two trajectories matched with it in the second device must also maintain the same relative order. If any two trajectories are found to be in reverse order, the candidate matching scheme is determined to violate physical laws. The cost of the corresponding matching position is set to an extremely large value, and the Hungarian algorithm is rerun until a globally optimal matching scheme that satisfies all order constraints is obtained.The optimal matching scheme determines the unique corresponding trajectory in the second device for each trajectory in the first device, which includes the matching pair of the target vehicle, meaning that the trajectory in the second device corresponding to the target vehicle is correctly identified.
[0074] If a consensus propagation algorithm is used, global optimization is achieved by constructing a Markov random field. First, each trajectory in the first device is considered a variable node, and the label space for each variable is all trajectories in the second device, i.e., the candidate set of second device trajectories that each trajectory in the first device might match. An energy function is defined to measure the overall cost of any matching scheme. This energy function consists of unary terms and pairwise terms. The unary term is the matching cost between a single trajectory and a candidate trajectory, calculated in the same way as the matching cost in graph matching algorithms, combining topological feature similarity and motion state consistency. The pairwise terms are used to impose logical consistency constraints on macroscopic traffic flow. For any two trajectories in the first device, if one trajectory is spatially preceding the other, the labels assigned to them in the second device must also maintain the same relative order; otherwise, a large penalty value is applied to the energy function. After defining the energy function, a belief propagation algorithm is used for iterative solution. In each iteration, each variable node sends a message to its neighboring nodes. The content of the message is based on the current node's local energy and the messages received from its neighbors, reflecting the current node's confidence in each label. Through multiple iterations, the message propagates and updates throughout the graph, gradually reducing the global energy and causing it to converge. After convergence, each variable node selects the label with the highest confidence as the final matching result. The resulting global matching scheme not only minimizes overall energy but also satisfies the physical laws governing all relative orders. This scheme also determines the corresponding trajectory of each trajectory in the first device in the second device, which includes the matching pairs of the target vehicle.
[0075] After obtaining the correct cross-camera matching pair using any of the above algorithms, the first and second vehicle trajectories of the target vehicle are spatiotemporally stitched together. First, based on the spatial distance between the two devices and the target average speed calculated from the first velocity vector of the target vehicle, the estimated travel time is obtained by dividing the spatial distance by the target average speed. Each timestamp of the second trajectory is shifted forward by this travel time to align the time axes of the two trajectories, compensating for the vehicle's travel time between the two cameras' fields of view. Then, using the rotation matrix and translation vector obtained during the camera calibration phase, the coordinate systems of the first and second 3D coordinates are transformed to the same world coordinate system, eliminating coordinate reference differences caused by different camera perspectives. For potential trajectory discontinuities at the stitching point, a linear interpolation method is used. This involves taking the coordinate values of the two points before and after the stitching point, calculating the coordinate values of the intermediate position according to the time ratio, and filling in the gaps to ensure a smooth trajectory transition. After the above processing, a complete motion path that is continuous in time, smooth in space, and conforms to the physical laws of vehicle relative order is formed—the continuous spatiotemporal trajectory of the target. Finally, this continuous spatiotemporal trajectory of the target is output as the target's global information.
[0076] In summary, the embodiments of this application have at least the following technical effects:
[0077] This application acquires a real-time image stream using pre-deployed image acquisition devices. The real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device. A first image corresponding to a target vehicle is selected from the first image stream, and the first image is analyzed to obtain first target information of the target vehicle. A target image is determined based on the second image stream, and the target image is analyzed to obtain second target information. The first target information and the second target information are fused and matched based on a cross-scene multimodal fusion strategy to obtain a preliminary matching result. If the preliminary matching result meets predetermined constraints, an alignment command is issued. Based on the alignment command, the first device and the second device undergo cross-viewpoint alignment processing to obtain global target information of the target vehicle. This invention solves the technical problems in the prior art where vehicle appearance recognition is affected by illumination, viewing angle, and occlusion, and trajectory association easily fails in dense traffic flow. By utilizing the dynamic topology between vehicles for feature fusion, it achieves the technical effect of improving the accuracy and robustness of cross-camera and cross-scene matching in complex traffic environments.
[0078] Example 2 is based on the same inventive concept as the vehicle cross-scene matching method based on the dynamic topology of the preceding vehicles in the previous example, such as... Figure 2 As shown, this application provides a vehicle cross-scenario matching system based on dynamic topology of front and rear vehicles. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0079] Image acquisition module 11 is used to acquire a real-time image stream through a pre-deployed image acquisition device, wherein the real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device; a first analysis module 12 is used to filter the first image corresponding to the target vehicle in the first image stream and analyze the first image to obtain the first target information of the target vehicle; a second analysis module 13 is used to determine the target image based on the second image stream and analyze the target image to obtain the second target information; a fusion matching module 14 is used to perform fusion matching on the first target information and the second target information based on a cross-scene multimodal fusion strategy to obtain a preliminary matching result; an instruction sending module 15 is used to issue an alignment instruction if the preliminary matching result meets a predetermined constraint; and a processing module 16 is used to perform cross-view alignment processing on the first device and the second device based on the alignment instruction to obtain the target global information of the target vehicle.
[0080] Furthermore, the system is also used to implement the following functions:
[0081] Read the deployment strategy and deploy the image acquisition device according to the deployment strategy; read the acquisition strategy and activate the image acquisition device to acquire data according to the acquisition strategy to obtain the real-time image stream.
[0082] Furthermore, the system is also used to implement the following functions:
[0083] The deployment strategy refers to deploying the image acquisition equipment at intervals of 50 meters to 200 meters along the monitored road section on the side of the road, with an installation height of 3 meters to 6 meters and a field of view covering 3 to 6 lanes. The image acquisition equipment refers to a monocular camera or a binocular camera.
[0084] Furthermore, the system is also used to implement the following functions:
[0085] The acquisition strategy refers to achieving data synchronization between cameras within the array through the PTP time synchronization protocol, with an acquisition frame rate of 25-30fps and a synchronization error of ≤10ms.
[0086] Furthermore, the system is also used to implement the following functions:
[0087] The first image is solved based on depth estimation technology to obtain a first spatiotemporal state sequence; the first spatiotemporal state sequence is added to the first target information; wherein, the first spatiotemporal state sequence includes a first three-dimensional coordinate, a first three-dimensional size, a first velocity vector and a first heading angle.
[0088] Furthermore, the system is also used to implement the following functions:
[0089] Based on the first image, a first target perception domain is determined with the target vehicle as the center, and a first star-shaped topology map of the first target perception domain is drawn; the first star-shaped topology map is encoded to obtain a first topological feature vector, and added to the first target information; wherein, any edge in the first star-shaped topology map has at least the identifiers of a first relative distance, a first relative speed, a first relative angle, a first TTC (time of collision), and a first lane relationship.
[0090] Furthermore, the system is also used to implement the following functions:
[0091] Calculate the spatial distance between the first device and the second device; combine the target average speed of the target vehicle and the spatial distance to filter the second image stream and obtain a candidate image set; extract any candidate image from the candidate image set and obtain any target information of the arbitrary candidate image; determine the target image with the highest similarity between the arbitrary target information and any information of the first target information.
[0092] Furthermore, the system is also used to implement the following functions:
[0093] The scene information is analyzed according to the cross-scene multimodal fusion strategy to obtain the matching confidence; the first target information and the second target information are fused and matched based on the matching confidence.
[0094] Furthermore, the system is also used to implement the following functions:
[0095] The first vehicle trajectory of the target vehicle collected by the first device is obtained; the second vehicle trajectory of the target vehicle collected by the second device is obtained; based on a predetermined algorithm and constrained by the logical consistency of macroscopic traffic flow, the target continuous spatiotemporal trajectory of the target vehicle is obtained, wherein the predetermined algorithm refers to a graph matching algorithm or a consistency propagation algorithm; the target continuous spatiotemporal trajectory is used as the target global information.
[0096] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0097] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A vehicle cross-scene matching method based on dynamic topology of front and rear vehicles, characterized in that, include: A real-time image stream is acquired through pre-deployed image acquisition devices, wherein the real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device; In the first image stream, filter the first image corresponding to the target vehicle, and analyze the first image to obtain the first target information of the target vehicle; The target image is determined based on the second image stream, and the target image is analyzed to obtain the second target information; The first target information and the second target information are fused and matched based on a cross-scenario multimodal fusion strategy to obtain a preliminary matching result; If the preliminary matching result meets the predetermined constraints, an alignment command is issued; Based on the alignment instructions, cross-view alignment processing is performed on the first device and the second device to obtain the target global information of the target vehicle.
2. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 1, characterized in that, Acquire real-time image streams through pre-deployed image acquisition devices, including: Read the deployment strategy and deploy the image acquisition device according to the deployment strategy; Read the acquisition strategy, activate the image acquisition device to acquire data according to the acquisition strategy, and obtain the real-time image stream.
3. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 2, characterized in that, The deployment strategy refers to deploying the image acquisition equipment at intervals of 50 meters to 200 meters along the monitored road section on the side of the road, with an installation height of 3 meters to 6 meters and a field of view covering 3 to 6 lanes. The image acquisition equipment refers to a monocular camera or a binocular camera.
4. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 2, characterized in that, The acquisition strategy refers to achieving data synchronization between cameras within the array through the PTP time synchronization protocol, with an acquisition frame rate of 25-30fps and a synchronization error of ≤10ms.
5. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 1, characterized in that, Filtering the first image corresponding to the target vehicle in the first image stream, and analyzing the first image to obtain the first target information of the target vehicle, including: The first image is solved based on depth estimation technology to obtain a first spatiotemporal state sequence; Add the first spatiotemporal state sequence to the first target information; The first spatiotemporal state sequence includes a first three-dimensional coordinate, a first three-dimensional dimension, a first velocity vector, and a first heading angle.
6. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 5, characterized in that, Add the first spatiotemporal state sequence to the first target information, followed by: Based on the first image, a first target perception domain is determined with the target vehicle as the center, and a first star-shaped topology map of the first target perception domain is drawn; The first star-shaped topology is encoded to obtain a first topological feature vector, which is then added to the first target information. In the first star-shaped topology map, any edge has at least the identification of a first relative distance, a first relative speed, a first relative angle, a first TTC (time of collision), and a first lane relationship.
7. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 1, characterized in that, Determining the target image based on the second image stream includes: Calculate the spatial distance between the first device and the second device; By combining the target average speed of the target vehicle and the spatial distance, the second image stream is filtered to obtain a candidate image set; Extract any candidate image from the candidate image set, and obtain any target information of the arbitrary candidate image; The target image is determined by taking the highest similarity between any target information and any information of the first target information.
8. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 1, characterized in that, The first target information and the second target information are fused and matched based on a cross-scenario multimodal fusion strategy to obtain a preliminary matching result, which includes: The scene information is analyzed according to the cross-scene multimodal fusion strategy to obtain the matching confidence. The first target information and the second target information are fused and matched based on the matching confidence level.
9. The vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in claim 1, characterized in that, Based on the alignment instructions, cross-view alignment processing is performed on the first device and the second device to obtain the target global information of the target vehicle, including: Obtain the first vehicle trajectory of the target vehicle collected by the first device; Acquire the second vehicle trajectory of the target vehicle collected by the second device; Based on a predetermined algorithm and constrained by the logical consistency of macroscopic traffic flow, the target continuous spatiotemporal trajectory of the target vehicle is obtained, wherein the predetermined algorithm refers to a graph matching algorithm or a consistency propagation algorithm. The continuous spatiotemporal trajectory of the target is used as the global information of the target.
10. A vehicle cross-scenario matching system based on dynamic topology of front and rear vehicles, characterized in that, The system is used to execute a vehicle cross-scene matching method based on dynamic topology of front and rear vehicles as described in any one of claims 1-9, the system comprising: An image acquisition module is used to acquire a real-time image stream through a pre-deployed image acquisition device, wherein the real-time image stream includes a first image stream acquired by a first device and a second image stream acquired by a second device; The first analysis module is used to filter the first image corresponding to the target vehicle in the first image stream and analyze the first image to obtain the first target information of the target vehicle; The second analysis module is used to determine the target image based on the second image stream and analyze the target image to obtain second target information; The fusion matching module is used to fuse and match the first target information and the second target information based on a cross-scenario multimodal fusion strategy to obtain a preliminary matching result; The instruction sending module is used to issue an alignment instruction if the preliminary matching result meets the predetermined constraints. The processing module is used to perform cross-view alignment processing on the first device and the second device based on the alignment instruction to obtain the target global information of the target vehicle.