Vehicle cooperative control method and system based on vehicle-road cooperative multi-target pedestrian recognition

By fusing multi-vehicle information and re-identifying pedestrians through roadside equipment, the trajectory prediction of the main vehicle is dynamically determined, which solves the problem of incomplete environmental perception among multiple vehicles and improves the safety and decision-making efficiency of the autonomous driving system.

CN121789451APending Publication Date: 2026-04-03SINO TRUK JINAN POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

When multiple vehicles approach pedestrian-dense areas simultaneously, existing autonomous driving systems lack effective cross-vehicle information sharing and fusion mechanisms, resulting in incomplete understanding of the environment by each vehicle, making it impossible to accurately identify and predict pedestrian targets, thus posing safety hazards.

Method used

Pedestrian image data is collected by multiple vehicles in a distributed manner and uploaded to roadside equipment. The roadside equipment performs information fusion, dynamically determines the main vehicle, and performs pedestrian re-identification and trajectory prediction. Collaborating vehicles make decisions based on unified prediction information.

Benefits of technology

It achieves global coverage perception of pedestrians on the road, improves the safety and reliability of multi-vehicle collaborative control, avoids resource waste and action conflicts, and ensures the consistency and accuracy of pedestrian behavior prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789451A_ABST
    Figure CN121789451A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent network connection automobile sensing, and particularly provides a vehicle cooperative control method and system based on vehicle-road cooperative multi-target pedestrian recognition, and the method comprises the steps: collecting pedestrian image data by a plurality of vehicles, extracting local sensing information containing appearance features and local position information, and uploading the local sensing information to roadside equipment; the roadside equipment performs pedestrian re-identification on the multi-source information, and dynamically determines a main vehicle for each pedestrian; the main vehicle performs motion state analysis and behavior prediction on the pedestrian to generate a prediction track; the main vehicle generates a driving state decision of the main vehicle based on the prediction information of the main vehicle and the prediction information of other main vehicles, and the cooperative vehicle generates a driving state decision of the main vehicle according to the prediction information of each main vehicle. According to the invention, the cooperative vehicles make decisions based on the unified prediction basis, the scene adaptability and decision-making efficiency are improved, the multi-vehicle action consistency is ensured, and the safety and reliability of vehicle cooperative driving in the multi-target pedestrian scene are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent connected vehicle perception, specifically to a vehicle cooperative control method and system based on vehicle-road cooperative multi-target pedestrian recognition. Background Technology

[0002] With the development of intelligent transportation and autonomous driving technologies, vehicle perception of the road environment, especially the stable identification and intention prediction of pedestrian targets, is crucial for ensuring driving safety and achieving high-level autonomous driving. However, pedestrian perception and control schemes in related autonomous driving systems generally adopt a single-vehicle independent perception and decision-making mode, relying on their own cameras, radar, and other sensors for environmental perception. When multiple vehicles approach a pedestrian-dense area (such as an intersection) simultaneously, each vehicle can only obtain partial pedestrian information from its own perspective. Due to the lack of an effective cross-vehicle information sharing and fusion mechanism, each vehicle's understanding of the environment is incomplete. For example, vehicle A may not see pedestrian P about to cross the road because of obstruction by a bus on its left, while vehicle B on its right can clearly observe P. Due to the lack of information sharing, vehicle A is unaware of this potential risk and may continue driving at its original speed, posing a safety hazard. Related vehicle-to-vehicle or vehicle-to-infrastructure communication schemes attempt to share raw or simply processed perception data. However, these solutions failed to effectively re-identify the perceived information from different vehicles, and could not determine whether the "pedestrian A" reported by vehicle A and the "pedestrian B" reported by vehicle B were the same actual pedestrian, resulting in information confusion and making it difficult to use directly for accurate decision-making. Summary of the Invention

[0003] To address the aforementioned issues, this invention provides a vehicle cooperative control method and system based on vehicle-road cooperative multi-target pedestrian recognition. This method enables cooperative vehicles to make decisions based on unified prediction criteria, improving scenario adaptability and decision-making efficiency, ensuring consistency of multi-vehicle actions, and enhancing the safety and reliability of vehicle cooperative driving in multi-target pedestrian scenarios.

[0004] The present invention provides a vehicle cooperative control method and system based on vehicle-road cooperative multi-target pedestrian recognition. Multiple vehicles collect pedestrian image data within their perception range, extract local perception information from the pedestrian image data, and upload the local perception information to roadside equipment. The local perception information includes at least pedestrian appearance features and local location information. The roadside equipment performs pedestrian re-identification by determining whether the perception information from different vehicles corresponds to the same actual pedestrian. For each collected pedestrian, a master vehicle is dynamically determined from the vehicles that can observe the pedestrian. Vehicles not responsible for any pedestrian are recorded as cooperative vehicles. Each master vehicle performs motion state analysis and behavior prediction for the pedestrian it is responsible for, generating prediction information containing the predicted trajectory of the pedestrian. The master vehicle generates its own driving state decision based on its own prediction information and the prediction information of other master vehicles, and the cooperative vehicles generate their own driving state decisions based on the prediction information of each master vehicle.

[0005] As can be seen from the above technical solutions, this application has the following advantages: (1) By collecting pedestrian image data in a distributed manner from multiple vehicles and uploading it to roadside equipment, the roadside equipment integrates multi-source local perception information, breaking the isolation of single-vehicle perception. As an information fusion center, the roadside equipment can aggregate the observation data of different vehicles, fill the perception blind spots of a single vehicle, form a global coverage perception of pedestrians on the road, solve the problem of incomplete environmental perception in the existing technology, and provide comprehensive and accurate basic data for multi-vehicle cooperative control; (2) For each pedestrian collected, the master vehicle is dynamically determined from the observed vehicles. The master vehicle is responsible for the core responsibility of analyzing pedestrian movement status and predicting behavior. The cooperating vehicles make decisions based on the prediction information of the master vehicle, which avoids redundant calculations by multiple vehicles, reduces resource waste, and improves decision response speed. Furthermore, as the responsible entity, the master vehicle can centrally process multi-view observation data, ensuring the consistency and accuracy of pedestrian behavior prediction and providing a reliable basis for collaborative control. (3) In both fully matched and blind spot scenarios, the cooperative vehicles use the pedestrian prediction trajectory generated by the master vehicle as the basis for decision-making, ensuring that the risk assessment and driving actions of multiple vehicles for the same pedestrian are consistent. In fully matched scenarios, the master vehicle and the cooperative vehicles make independent decisions based on unified prediction information to avoid action conflicts; in blind spot scenarios, the cooperative vehicles that have not observed the pedestrian receive the prediction information from the master vehicle to predict and avoid potential risks, thereby improving the safety and coordination of multi-vehicle cooperative driving. Attached Figure Description

[0006] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0007] Figure 1 This is a schematic flowchart of a vehicle cooperative control method based on multi-target pedestrian recognition in vehicle-road cooperative manner, provided in an embodiment of the present invention.

[0008] Figure 2 This is a schematic block diagram of a vehicle cooperative control system based on multi-target pedestrian recognition in vehicle-road cooperative manner, provided as an embodiment of the present invention. Detailed Implementation

[0009] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0010] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this application and in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0011] Figure 1 This is a flowchart illustrating a vehicle cooperative control method based on multi-target pedestrian recognition in vehicle-road cooperative system, provided by an embodiment of the present invention. The order of the steps in this flowchart can be changed or some can be omitted depending on different requirements.

[0012] S1, multiple vehicles collect pedestrian image data within their perception range, extract local perception information from the pedestrian image data, and upload the local perception information to the roadside equipment; the local perception information includes at least pedestrian appearance features and local location information.

[0013] S2, the roadside equipment performs pedestrian re-identification by determining whether the perception information from different vehicles corresponds to the same actual pedestrian. For each pedestrian that is collected, a master vehicle is dynamically determined from the vehicles that can observe the pedestrian; vehicles that are not responsible for any pedestrian are recorded as cooperative vehicles.

[0014] S3 involves each master vehicle analyzing the motion state and predicting the behavior of the pedestrians it is responsible for, generating prediction information containing the predicted trajectory of that pedestrian.

[0015] S4, the master vehicle generates its own driving state decision based on its own prediction information and the prediction information of other master vehicles, and the cooperating vehicle generates its own driving state decision based on the prediction information of each master vehicle.

[0016] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process in this embodiment, another vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition is provided, which includes the following steps.

[0017] S101, vehicles collect pedestrian image data and extract local perception information, then upload the local perception information to roadside equipment.

[0018] S101.1, Multiple vehicles collect pedestrian image data within their perception range respectively.

[0019] The vehicle includes at least one image acquisition module, one positioning and timing module, and one onboard computing unit. The image acquisition module is typically mounted inside the vehicle's windshield or on the roof, with its optical axis substantially parallel to the vehicle's longitudinal axis. It is pre-calibrated to obtain accurate intrinsic and extrinsic camera parameters. The intrinsic camera parameters include focal length, optical center, and distortion coefficients, while the extrinsic parameters include the mounting position and angle relative to the vehicle's coordinate system.

[0020] Data acquisition is triggered by the vehicle's own perception cycle or external collaborative events. In one implementation, the method automatically initiates when a vehicle enters a specific collaborative zone broadcast by roadside equipment. To ensure temporal comparability of multi-vehicle data, each vehicle relies on a high-precision timing module to assign a unified timestamp to each frame of acquired image data, with time synchronization accuracy reaching the millisecond level. The specific collaborative zone can be the geofence area of ​​an intersection.

[0021] Pedestrian image data refers to raw or pre-processed digital images containing potential pedestrian targets from the sides, front, and intersections of the road. After acquisition, the onboard computing unit can perform necessary pre-processing operations on the raw images, including geometric correction and format conversion.

[0022] Because vehicles differ in their position, direction of travel, and speed on the road, their perception range forms a spatial observation network with multiple perspectives, overlapping, and non-overlapping areas. In overlapping areas, multiple vehicles may observe the same group of pedestrians from different angles, providing multi-source data for subsequent roadside recognition. In non-overlapping areas, a perception blind spot created by a vehicle due to obstruction by buildings or large vehicles may be covered by the sensors of another vehicle in a different location. For example, a vehicle traveling on the north side of an intersection may not see a pedestrian crossing from the south side, obscured by a building on the east side, while a vehicle approaching from the west side may observe them.

[0023] S101.2 Extract local perception information from pedestrian image data.

[0024] S101.21 The acquired real-time image frames are processed by the target detection algorithm to locate all pedestrian targets in the image, and a bounding box and corresponding detection confidence score are output for each target.

[0025] The vehicle-mounted computing unit loads an object detection model, which can be YOLOv5, YOLOX, or other lightweight variants. It takes a preprocessed single-frame image as input, performs forward inference, and outputs bounding boxes and detection confidence scores.

[0026] Specifically, for each pedestrian target identified in the image, the bounding box coordinates in the image pixel coordinate system are output, represented as... ,in The coordinates of the top left corner of the bounding box. The coordinates are the bottom right corner.

[0027] The detection confidence score is a value between 0 and 1, representing how certain the model is that a pedestrian target exists within the bounding box. A higher confidence score indicates a more reliable detection result.

[0028] S101.22 For each detected pedestrian target, the image region within its bounding box is extracted as the region of interest, and the region is input into a pre-trained lightweight convolutional neural network for multi-dimensional feature extraction; the extracted multi-dimensional features include at least clothing category features, color distribution features, and accessory features.

[0029] Based on the bounding box coordinates, a corresponding rectangular region is cropped from the original image as the region of interest (ROI) for the pedestrian. This ROI is then input into a pre-trained lightweight convolutional neural network, which is designed as a multi-task or cascaded structure capable of extracting multiple appearance features in parallel or sequentially from a single input.

[0030] For clothing category features, a classification branch of the network outputs a feature vector representing the type of clothing worn by the pedestrian, specifically their top and bottom garments. For color distribution features, a feature layer or a separate processing branch of the network extracts and encodes the color information of the pedestrian's main body region. For example, the image is converted to the HSV or Lab color space, its primary color histogram is calculated, and then compressed into a low-dimensional vector. For accessory features, an object detection or attention branch of the network identifies and encodes the features of salient items carried by the pedestrian, such as backpacks, handbags, hats, etc. If there are no obvious accessories, a feature vector representing "none" or close to zero is output.

[0031] The network architecture includes a shared feature extraction backbone, a multi-task feature extraction head, and a feature fusion and output layer.

[0032] The shared feature extraction backbone uses a lightweight convolutional neural network that has been deeply compressed and optimized, such as MobileNetV3-Small, ShuffleNetV2, or EfficientNet-Lite. This backbone network receives a preprocessed region of interest (ROI) image as input and outputs a shared intermediate feature map F_shared. F_shared encodes general visual information from the image.

[0033] Following F_shared, the network connects three independent task-specific branches in parallel, each responsible for extracting features in one dimension.

[0034] The clothing category header consists of 1-2 convolutional layers, a global average pooling layer, and a fully connected layer, outputting a C_cls-dimensional feature vector f_attire, where each dimension corresponds to the response intensity of a predefined clothing category. During the pre-training phase, supervised classification training is performed using data labeled with clothing categories to learn discriminative features related to clothing.

[0035] The color distribution head first maps F_shared to a specific number of channels using a convolutional layer, then performs global average pooling, followed by a small fully connected layer. The output is a C_color dimensional vector f_color, which encodes the statistical distribution of the image in a specific color space. This can be a color representation learned by the network, or a distribution vector obtained by calculating similarity with a pre-defined color vocabulary. During training, color histograms are used as soft targets for regression training, or color cluster centers are used as classification labels for classification training.

[0036] The accessory detection head consists of an upsampling layer and several convolutional layers, outputting a spatial attention map or a small bounding box prediction. Specifically, it outputs a C_acc-dimensional vector f_accessory, indicating the presence of backpacks, tote bags, hats, etc. If no accessories are found, it outputs a vector close to zero. Training is performed using data with accessory bounding boxes or presence labels, employing either binary cross-entropy loss or focus loss.

[0037] The feature fusion and output layer ultimately concatenates the output vectors of the three task heads: F_final=Concat(f_attire,f_color,f_accessory). The lightweight feature vector F_final integrates information from the three dimensions of clothing, color, and accessories.

[0038] The training process includes: collecting a general pedestrian image dataset and performing multi-task annotation, including: pedestrian ID (for re-identification), clothing category label, color attribute label, and labels for the presence and location of accessories. A multi-task joint loss function L_total is used.

[0039] in: For pedestrian identification loss, such as cross-entropy loss or triplet loss; These are the loss functions for clothing classification, color attribute classification / regression, and accessory detection, respectively. These are the weighting coefficients for each task, used to balance the learning speed of different tasks.

[0040] The entire network is trained end-to-end using the gradient descent algorithm.

[0041] S101.23, Based on the pixel coordinates of the pedestrian target's bounding box in the image and the vehicle's inherent parameters, calculate the local position information of the pedestrian relative to the vehicle's coordinate system; the local position information includes at least the relative lateral distance, relative longitudinal distance, and azimuth angle between the pedestrian and the vehicle.

[0042] Specifically, select the pixel coordinates of the bottom midpoint of the pedestrian bounding box. This serves as an approximate projection point of its feet contacting the ground. An inverse perspective transformation is performed based on the vehicle's inherent parameters, including the camera intrinsic matrix and camera extrinsic parameters.

[0043] Assuming the ground is a plane, pixel coordinates Transform to vehicle coordinate system Below, the longitudinal distance along the direction of the vehicle's front. Lateral distance perpendicular to the direction of the front of the vehicle The azimuth angle represents the direction angle of the line connecting the pedestrian and the vehicle. . This refers to the pedestrian's local position relative to the vehicle. Among these, The mounting height of the camera relative to the vehicle's coordinate system. For camera focal length, This is the camera's principal point.

[0044] S101.24, the extracted multi-dimensional features are concatenated to form a lightweight feature vector, and the lightweight feature vector is associated and encapsulated with local location information, target detection timestamp, and temporary local identifier of the pedestrian target to form a data packet of the local perception information.

[0045] The clothing category feature vector, color distribution feature vector, and accessory feature vector extracted in S101.22 are concatenated to form a comprehensive lightweight feature vector F_app. A temporary local identifier Local_ID, unique within this vehicle and this frame, is assigned to the pedestrian target.

[0046] Create a structured data packet that contains at least the following fields: Local_ID: Temporary local identifier; F_app: Lightweight feature vector; Pos_local: Local location information (X,Y,α); TimeStamp: Target detection timestamp; Confidence: Confidence in object detection; Vehicle_ID: The unique identifier for this vehicle.

[0047] S101.3 Upload local sensing information to roadside equipment.

[0048] The vehicle communicates with roadside equipment via an integrated vehicle-to-everything (V2X) communication module. Preferably, cellular network-based vehicle wireless communication technology is used. The onboard computing unit encapsulates the local perception information data packets generated in S101.24 and transmits the encapsulated data packets to the roadside equipment.

[0049] In S102, roadside equipment re-identifies pedestrians and assigns a master vehicle to the collected pedestrians based on the re-identification results.

[0050] S102.1, the roadside equipment performs pedestrian re-identification on the received multi-source local sensing information to determine whether sensing information from different vehicles corresponds to the same pedestrian.

[0051] S102.11, the local location information in the data packet is uniformly converted to the global coordinate system maintained by the roadside equipment to obtain the global location estimate of each observed pedestrian.

[0052] The roadside equipment maintains a global coordinate system. For each data packet, it obtains the global positioning and attitude of the sending vehicle at the time of data acquisition from the data packet header. It combines the local position information recorded in the data packet relative to the coordinate system of the sending vehicle with the global pose of the sending vehicle and calculates the approximate coordinates of the pedestrian in the global coordinate system through spatial geometric transformation, thereby obtaining the global position estimate of each observed pedestrian.

[0053] S102.12 decodes the lightweight feature vector to obtain the clothing category feature vector, color distribution feature vector, and accessory feature vector.

[0054] The roadside equipment performs a reverse operation on the lightweight feature vector in each data packet, including splitting F_final into three independent sub-vectors according to the known dimensions of each feature vector, to obtain the clothing category feature vector, color distribution feature vector, and accessory feature vector.

[0055] S102.13, For all pedestrian observation data received at the current moment, for any two pedestrian observation data from different vehicles, perform matching operations a) to c).

[0056] a) Retrieve three parallel CSPNet residual network branches. The first, second, and third branches process the clothing category feature vector, color distribution feature vector, and accessory feature vector, respectively. When each branch performs forward propagation, it introduces the feature maps of the other two branches in the target intermediate layer through cross-branch connections. The introduced feature maps are concatenated or element-wise added with the residual edge output of the current branch to output their respective fused feature vectors. Calculate the Euclidean distance between the corresponding fused feature vectors and map the similarity score of the corresponding dimension based on this distance.

[0057] The roadside equipment loads a pre-trained multi-branch CSPNet residual network, which contains three branches with the same structure but independent weights, specifically for handling clothing, color, and accessory features.

[0058] For two observation data points, Obs_A and Obs_B, to be matched, their decoded feature vectors are input into the corresponding branches. However, when the first branch (clothing branch) processes the clothing category feature vector, in addition to its own convolution calculation, it also introduces feature maps generated by the second branch (color branch) and the third branch (accessory branch) when processing the color distribution feature vector and accessory feature vector in their respective intermediate layers. These introduced feature maps are then concatenated or element-wise added to the residual edge output of the current layer of the first branch. The second and third branches perform the same operation, introducing feature maps from the other two branches respectively. After this cross-branch fusion process, each branch ultimately outputs an enhanced feature vector that integrates multi-dimensional information.

[0059] Calculate the Euclidean distance between the fused feature vectors of Obs_A and Obs_B in three dimensions, and then map each distance to a similarity score between 0 and 1 using a trainable Sigmoid function.

[0060] The multi-branch CSPNet residual network consists of an input layer, three parallel CSPNet residual branches, a cross-branch connection and fusion mechanism, and an output layer.

[0061] The input layer accepts three independent input vectors, corresponding to the decoded clothing feature vector f_attire, color feature vector f_color, and accessory feature vector f_accessory, respectively. Each vector is first passed through a fully connected layer or a one-dimensional convolutional layer, mapped to a unified feature dimension, and reshaped into a two-dimensional feature map.

[0062] Each of the three parallel CSPNet residual branches is based on the CSPNet architecture, which effectively mitigates gradient vanishing, enhances feature reuse, and reduces computation by splitting the feature map into two parts and fusing them at different stages. The structure of each branch is as follows: Stage splitting: The input feature map is evenly split into two parts, Part1 and Part2; Residual path: Part 1 undergoes deep feature transformation through a series of residual blocks; Direct path: Part 2 goes through a simple convolutional layer or is retained directly; Feature fusion: The output of the residual path and the output of the direct path are concatenated along the channel dimension to form the output feature map of this stage.

[0063] Each branch consists of several such CSP stages stacked together, progressively extracting and refining features of the corresponding dimension.

[0064] Cross-branch connection and fusion mechanism refers to performing cross-branch information exchange at a specific intermediate layer in the three parallel branches. This specific intermediate layer can be at the end of each CSP stage. For the currently processed target branch, the system introduces intermediate feature maps from the other two branches at the same network depth through a cross-branch connection layer. The cross-branch connection layer is a 1x1 convolutional layer. The feature maps from the other two branches are then fused with the output feature map of the current layer of the target branch. The fusion method can be channel concatenation or element-wise addition.

[0065] After multiple CSP stages with cross-branch fusion, the output of each branch is mapped to a low-dimensional feature embedding space through a global average pooling layer and a final fully connected layer. Ultimately, the three branches output fused clothing feature vectors, fused color feature vectors, and fused accessory feature vectors, respectively.

[0066] Using a large-scale pedestrian re-identification dataset or a self-built dataset, each sample contains multiple images of the same pedestrian from different cameras, at different times, and from different angles. For each image, a pre-trained feature extractor is used to extract the [f_attire, f_color, f_accessory] vector, which is then used as the input to this network.

[0067] During training, a combined loss function is used, including triplet loss, identity classification loss, and total loss.

[0068] For each "anchor" sample, select a positive sample (from the same pedestrian) and a negative sample (from different pedestrians). Calculate the distance between the anchor and the positive and negative samples on the three fused feature vectors, and apply the triplet loss to each:

[0069] in, This represents the triplet loss for the i-th feature dimension. Let i be the feature vector of the anchor sample in the i-th dimension. Let i be the feature vector of the positive sample in the i-th dimension. Let i be the feature vector of the negative sample in the i-th dimension. Let be the Euclidean distance function. This is a preset positive boundary value used to control the distance difference between positive and negative sample pairs.

[0070] Add a classification header to the output of each branch to predict the pedestrian's identity ID, using standard cross-entropy loss. .

[0071] Total loss:

[0072] in, For the overall joint loss, and These are hyperparameters used to balance the weights of the learning loss and the classification loss, respectively. C={attire,color,accessory} represents the set of feature dimensions.

[0073] Training is conducted in stages. In the first stage, each branch is pre-trained independently, with cross-branch connections temporarily disabled. Each of the three branches is pre-trained independently using its own input features and identity classification loss to initially learn feature representations for each dimension. In the second stage, joint fine-tuning is performed, enabling cross-branch connections and using the overall joint loss to fine-tune the entire network end-to-end.

[0074] b) Calculate the physical distance based on the global position estimate after the transformation of the two observation data, and determine whether the physical distance is within the preset movement speed range of the pedestrian based on the timestamp difference between the two data. Map the spatial consistency score according to the judgment result.

[0075] The physical distance between the two observations is calculated based on the global position estimates obtained from S102.11, and the time difference Δt is calculated by obtaining the timestamps of the two data packets.

[0076] Set a reasonable maximum speed for a pedestrian, and calculate the maximum distance the pedestrian can move within a time difference Δt = maximum speed * time difference Δt.

[0077] If the calculated physical distance is less than or equal to the maximum distance a pedestrian may move, the two observations are considered to be consistent in space and time, and a high spatial consistency score is assigned. Otherwise, they are considered to be inconsistent in space and time, and the high spatial consistency score is reduced to a degree proportional to the difference between the physical distance and the maximum distance a pedestrian may move, thus obtaining the spatial consistency score.

[0078] c) Determine whether two observations correspond to the same actual pedestrian based on spatial consistency score and similarity scores in each dimension.

[0079] The spatial consistency score must be higher than the basic threshold. If the similarity scores of the three appearance dimensions are not lower than their respective judgment thresholds, then Obs_A and Obs_B are determined to correspond to the same actual pedestrian. Alternatively, dynamic weights are assigned to the three appearance similarity scores, and a weighted total score is calculated. If the weighted total score is not lower than the threshold, then Obs_A and Obs_B are determined to correspond to the same actual pedestrian.

[0080] Furthermore, based on the pedestrian re-identification results, the current scene type is determined, including: maintaining a valid observer set for each collected pedestrian, forming a mapping relationship between pedestrians and vehicles that observed them; determining the relevant vehicle set and the entire set of observed pedestrians within the current collaborative area; if every pedestrian in the entire pedestrian set can be observed by all vehicles in the relevant vehicle set, it is determined to be a perfectly matched scene. If at least one pedestrian in the entire pedestrian set cannot be observed by at least one vehicle in the relevant vehicle set, it is determined to be a scene with a blind spot.

[0081] In other words, within this collaborative area, if every vehicle can observe every single pedestrian, and no vehicle has a blind spot in its perception of any pedestrian, it is determined to be a perfectly matched scenario. If at least one pedestrian cannot be observed by any one or more vehicles in the relevant vehicle set, this indicates that there is incomplete overlap in the perception coverage between vehicles, resulting in a blind spot scenario.

[0082] For example, suppose at an intersection, the cooperative vehicles are V_region={vehicle A, vehicle B}, and the observed pedestrians are P_total={pedestrian X, pedestrian Y}. If the effective observer set O(pedestrian X)={vehicle A, vehicle B} and the effective observer set O(pedestrian Y)={vehicle A, vehicle B}, then it is determined to be a perfectly matched scene. If the effective observer set O(pedestrian X)={vehicle A, vehicle B}, but the effective observer set O(pedestrian Y)={vehicle A} (vehicle B does not see Y due to occlusion), then it is determined to be a scene with a blind spot, where pedestrian Y constitutes a perception blind spot for vehicle B.

[0083] S102.2 For each pedestrian collected, a master vehicle is dynamically determined from the vehicles that can observe the pedestrian.

[0084] S102.21, the roadside equipment maintains a list of vehicles that can currently observe each pedestrian for each pedestrian that is collected.

[0085] For each pedestrian P in the global pedestrian identity list j Roadside equipment from its associated set of valid observers O(P) j From the current time window, extract all vehicles that are consistently reporting pedestrian data, forming a candidate vehicle list (Candidates(P)). j For example, Candidates(P) j ={V1,V2,...,V N}

[0086] S102.22 Calculate the suitability score for each candidate device based on at least one of the following factors: the relative distance between each candidate vehicle and the pedestrian, the quality of its reported observation data, and the stability of its historical observations.

[0087] The quality of the observation data is characterized by the target detection confidence and the observation perspective of the candidate device on the pedestrian; the stability of the historical observation is characterized by the number of consecutive frames or frequency in which the candidate device has successfully reported the pedestrian data within a preset time period in the past.

[0088] The suitability score integrates evaluation factors from multiple dimensions to quantify the suitability of each candidate vehicle (V). k As P j The scoring function for the quality of the main vehicle is as follows:

[0089] In the formula, For candidate vehicle V k For pedestrian P j Suitability score, These are the weighting coefficients for the distance factor, data quality factor, and stability factor, respectively, which can be preset or adaptively adjusted according to the scenario.

[0090] This is the distance factor function, with vehicle V as the input. k With pedestrian P j relative distance between This distance is calculated based on the latest global positions of both objects. The function is monotonically decreasing; the closer the distance, the higher the score. One possible implementation is... For a linear mapping function, it is expressed as:

[0091] In the formula, This is the preset maximum effective election distance; vehicles exceeding this distance will not be considered.

[0092] This is the data quality factor function. The input is vehicle V. k Reported about P j Observational data quality . It consists of two factors: the detection confidence level (conf). k and observation perspective rating view k Detect confidence level conf k Target detection confidence from data packets, viewpoint score k According to vehicle V k With pedestrian P j The relative azimuth angle is calculated by defining the side as the optimal viewing angle, which receives the highest score, while the front or rear angles receive lower scores. The calculation formula is as follows:

[0093] In the formula, The angle between the direction the vehicle is pointing towards the pedestrian and the direction the pedestrian is directly to the side.

[0094] Overall quality Data quality factor function The sigmoid function is used to increase the quality and thus the score.

[0095] This is a historical observation stability factor function, with vehicle V as the input. k In recent times, regarding P j Observational stability measure . It can be characterized by calculating the proportion of successfully reported frames to the total number of frames (reporting frequency) or the number of consecutively successfully reported frames (consecutive frames). For a monotonically increasing function, the higher the stability, the higher the score, expressed as:

[0096] S102.23, the candidate vehicle with the highest suitability score is determined as the pedestrian's current primary vehicle.

[0097] For pedestrian P j The roadside equipment selects the vehicle with the highest suitability score from all candidate vehicles as its current primary vehicle. The roadside equipment generates a primary vehicle assignment command, which must include at least: <Pedestrian Global ID:P> j Master ID: Master(P) j The command, displaying the latest pedestrian status, is sent to the selected Master vehicle (P) via the vehicle-to-infrastructure (V2I) link. j ).

[0098] For example, suppose there is pedestrian P and candidate vehicles A and B. A =5m,Q A =0.9,S A =0.8; D B =8m,Q B =0.95,S B =1.0. Let the weight w D =0.5,w Q =0.3,w S =0.2, and all factor functions are linearly normalized. Therefore, H A =0.5*(1-5 / 20)+0.3*0.9+0.2*0.8=0.375+0.27+0.16=0.805, H B =0.5*(1-8 / 20)+0.3*0.95+0.2*1.0=0.3+0.285+0.2=0.785. Therefore, vehicle A is chosen as the main vehicle, even though its data quality is slightly lower, its advantage of being closer is greater.

[0099] Through the above multi-factor weighted election mechanism, the system can dynamically and adaptively select the vehicle that is spatially closest, has the best data quality, and has the most stable observation for each pedestrian as its analysis leader, thereby optimizing the allocation of perception resources and the source quality of decision data.

[0100] In this embodiment, a master vehicle is assigned to each pedestrian whose trajectory is captured. The master vehicle then predicts the pedestrian's trajectory. It's important to note that "captured pedestrians" refers to the aggregate of all pedestrians identified by all vehicles; as long as at least one vehicle identifies a pedestrian, that pedestrian is considered captured. Since each captured pedestrian has a master vehicle for trajectory prediction, the prediction information is shared among all vehicles requiring coordination within the collaborative area. This means that regardless of whether it's a perfectly matched scenario or a scenario with blind spots, all vehicles can obtain the prediction information for all captured pedestrians, and then each vehicle can assess the safety level of each pedestrian and make a decision. Therefore, even if a vehicle has blind spots, it can still drive safely through coordination with other vehicles.

[0101] After the Roadside Unit (RSU) assigns a master vehicle to each pedestrian, it generates a master vehicle assignment command. This command has the format: <Pedestrian Global ID: P_j, Master Vehicle ID: Master(P_j), Latest Status>. This command is sent point-to-point to the designated master vehicle via the vehicle-to-infrastructure (V2I) communication link. When a vehicle receives the master vehicle assignment command from the RSU, it parses the command content and learns that it has been assigned as the master vehicle for the corresponding pedestrian. The vehicle's internal state manager records this assignment relationship, for example, by maintaining a "primary responsible pedestrian list" in memory: {Responsible Pedestrians: P_j, ...}. A vehicle may be assigned as the master vehicle for multiple pedestrians by the RSU, and the vehicle will receive multiple master vehicle assignment commands (each for a different pedestrian), thus learning that it is responsible for multiple pedestrians. Its "primary responsible pedestrian list" will contain multiple entries, such as {Responsible Pedestrians: P_j1, P_j2, ...}.

[0102] After the roadside equipment completes the primary vehicle assignment for all pedestrians, all vehicles within the cooperative area that have not received any primary vehicle assignment commands are automatically defined as cooperative vehicles. If a vehicle does not receive any primary vehicle assignment command for itself within a preset time window, its logic judgment unit marks its status as "cooperative vehicle". To ensure reliability, the roadside equipment can broadcast a "role confirmation broadcast" after completing all primary vehicle assignments. This broadcast message contains a list: <Pedestrian P_1 Primary Vehicle: V_a, Pedestrian P_2 Primary Vehicle: V_b,..., Cooperative Vehicle: [V_c, V_d,...]>. After receiving this broadcast, all vehicles can cross-check their roles: if their own ID appears in the primary vehicle field of a pedestrian, they are confirmed as the primary vehicle for that pedestrian; if their own ID only appears in the "cooperative vehicle" list, or does not appear in the broadcast, they are confirmed as a cooperative vehicle.

[0103] After assigning the primary vehicle and notifying all parties, the roadside equipment broadcasts a global pedestrian list. All vehicles, including the primary vehicle and cooperating vehicles, receive this list, ensuring all vehicles are aware of which pedestrians require attention within the current cooperative area. After predicting the trajectories of the pedestrians under its responsibility, the primary vehicle encapsulates the data according to an agreed-upon format (including predicted trajectory, timestamp, pedestrian ID, and primary vehicle ID). This prediction information is broadcast to all other vehicles within the cooperative area via vehicle-to-vehicle (V2V) or via roadside equipment (V2I2V). Upon receiving prediction packets from each primary vehicle, the cooperating vehicles: match the global pedestrian ID field in the packet with the previously received "global pedestrian list" from the roadside equipment to ensure information validity. The cooperating vehicle clearly knows the source of each pedestrian prediction information because the packet contains the primary vehicle ID field; the cooperating vehicle then aggregates all received prediction information to form a complete local "pedestrian prediction information database" for localized risk assessment and decision-making.

[0104] S103, Pedestrian trajectory prediction.

[0105] Each master vehicle performs motion state analysis and behavior prediction on the pedestrians it is responsible for, and generates prediction information containing the predicted trajectory of the pedestrian, specifically including the following steps S103.1 to S103.3.

[0106] S103.1 The main vehicle acquires multi-view observation data of itself and other vehicles on the same pedestrian. Based on the pedestrian's torso target point as the rotation center, the pedestrian image is rotated and transformed to unify the multi-view observation data to a standard analysis perspective.

[0107] The main vehicle receives asynchronous observation data packets from itself and other vehicles regarding the same pedestrian via the V2X communication protocol. Each data packet contains at least a timestamp, the device's global coordinates, a pedestrian bounding box image, and corresponding preliminary skeletal keypoint data. The main device uses spatiotemporal alignment and appearance feature matching algorithms to confirm that these data packets point to the same pedestrian target.

[0108] In each frame of a pedestrian image, 2D skeletal keypoints of the human body are extracted using a pre-trained human pose estimation model, such as OpenPose or HRNet. The torso target point is defined as the midpoint of the line connecting the neck keypoint and the hip keypoint. This point is located on the approximate axis of symmetry of the human body and its position relative to the human body is relatively stable during walking, making it suitable as a reference for rotational transformations.

[0109] The standard analysis viewpoint is defined as the perspective of the pedestrian facing the observer. For each frame from any other device, the vector difference between the current torso target point and its correct position under the standard viewpoint is calculated. Based on this vector difference, a rotation transformation matrix is ​​calculated. This rotation transformation is applied to all pixels and extracted skeletal keypoints in the frame to generate a pedestrian image observed from the standard viewpoint and its corresponding keypoint coordinates.

[0110] The data from all other devices, after being unified in perspective, are interpolated and aligned according to timestamps, and then fused into multi-frame time-series data consistent with the pedestrian's perspective.

[0111] S103.2 Extract the gait temporal features of the pedestrian from the unified data. The gait temporal features include at least the changes in the angle between the leg bone points, the step frequency, and the step length.

[0112] Leg skeletal point angle variation extraction: In each frame, points related to leg movement are selected from the unified skeletal keypoints, such as: left hip (LH), left knee (LK), left ankle (LA); right hip (RH), right knee (RK), right ankle (RA). Key angles are calculated, such as the angle between the thigh and lower leg, which directly reflects the degree of knee flexion. These angle values ​​are arranged in chronological order to form a temporal sequence of angle variations. This sequence includes information such as gait cycle, gait speed, and gait anomalies.

[0113] Step frequency extraction: The above-mentioned leg angle change sequence or foot height sequence calculated by skeletal points is periodically analyzed using autocorrelation function or spectrum analysis methods to find the fundamental frequency component in the spectrum, which corresponds to the pedestrian's step frequency.

[0114] Stride length estimation: The actual spatial scale is estimated by using the ratio of the pedestrian's height pixel value in the image to the actual average height. In consecutive frames, the stride length can be estimated by tracking the horizontal displacement of foot keypoints under a unified viewpoint, combined with the above spatial scale conversion, and considering the ground plane assumption. Finally, a stride length time series is obtained.

[0115] S103.3, Based on the extracted gait temporal features, the predicted future position of the pedestrian is calculated using a temporal prediction model, and a predicted trajectory is generated. The temporal prediction model is expressed as follows:

[0116]

[0117] In the formula, For the predicted step size of the next time step, An adaptive weighting coefficient that is positively correlated with the pedestrian's current speed. These are the model coefficients. This represents the change in hip angle at the current moment. This represents the predicted positions of the pedestrian at time t and time t+1 in the global coordinate system. The predicted direction of motion is obtained by a moving average of the historical position sequence, for example, by calculating the average direction angle of a vector formed by the past m position points.

[0118] in, This constitutes an adaptive weighted historical step-size inertia term. For step frequency driven terms, It is a posture adjustment type.

[0119] Obtain the predicted step size and predicted direction Then, using the displacement decomposition formula under the uniform circular motion model, the position at time t+1 is calculated. .

[0120] S104, Collaborative Control Process.

[0121] S104.1 The master vehicle generates its own driving status decision based on its own prediction information and the prediction information of other master vehicles.

[0122] S104.11, convert the predicted trajectories of pedestrians under the responsibility of other master vehicles to the vehicle's coordinate system to obtain the converted local predicted trajectories.

[0123] Using the received transmitted vehicle pose and the current global pose of the vehicle, the rigid body transformation matrix from the transmitting vehicle coordinate system to the vehicle coordinate system is calculated. This matrix includes rotation and translation components. For each global coordinate point in the pedestrian's predicted trajectory, the transformation matrix is ​​applied to calculate its corresponding coordinates in the vehicle coordinate system, thus obtaining the transformed predicted trajectory of the pedestrian in the vehicle coordinate system.

[0124] S104.12 For each pedestrian collected, calculate the predicted minimum distance and safe braking margin time based on their predicted trajectory.

[0125] It should be noted that for pedestrians under the responsibility of the vehicle itself, the predicted trajectory is the trajectory predicted by the vehicle itself, while for pedestrians under the responsibility of other vehicles, the predicted trajectory is the local predicted trajectory obtained after conversion in step S104.11.

[0126] Step 1.1) From the pedestrian's predicted trajectory, select the point that is closest to the vehicle's expected path space within a fixed future time window as the nearest interaction point, and record the predicted location and arrival time of that point.

[0127] Input data includes predicted pedestrian trajectories. , where T is the fixed prediction time window and the vehicle's expected driving path.

[0128] Iterate through all predicted locations of the pedestrian within the predicted time window. For each predicted location... Calculate the shortest spatial distance between the predicted pedestrian point and the expected path of the vehicle, and select the pedestrian prediction point that minimizes this distance as the nearest interaction point.

[0129] Output the predicted position CPI_Pos of the nearest interaction point in the global coordinate system and the predicted time CPI_Time required for the pedestrian to move to that point from the current time t.

[0130] Step 1.2) Calculate the predicted minimum distance and safe braking margin time based on the vehicle's current position, current speed, and preset deceleration capability.

[0131] a) Predict the minimum distance.

[0132] The predicted minimum distance is the Euclidean distance between the vehicle's current position and the predicted position when it travels at a constant speed to the position at the arrival time.

[0133] b) Safety braking margin time.

[0134] The braking distance of the vehicle is calculated based on its current speed and maximum safe deceleration: (current speed) 2 / (2 * maximum safe deceleration).

[0135] Calculate the time required for the vehicle to come to a complete stop from the start of braking = current speed / vehicle braking stopping distance.

[0136] The distance the pedestrian moves within this time is defined by the sum of the distance the vehicle stops, the distance the pedestrian moves within this time, and the preset static safety margin, starting from the current position of the vehicle and extending backward along the current driving direction. This spatial area constitutes the safety envelope. The predicted time required for the pedestrian's trajectory to first enter the safety envelope from the current position, moving at its current speed and direction, is called the safety braking margin time.

[0137] Consider the pedestrian as a point moving at its current position, speed, and direction. Calculate the predicted time required for this point's trajectory to first enter the aforementioned safety envelope, starting from the current moment; this time is the safety braking margin time.

[0138] S104.13, based on the predicted minimum distance and safe braking margin time corresponding to the pedestrian, a risk level is assigned to the pedestrian by the judgment rule.

[0139] Risk levels include dangerous, warning, and safe.

[0140] A vehicle is classified as hazardous if it meets any of the following conditions: the predicted minimum distance is less than the hazardous distance threshold, or the safe braking margin time is less than the emergency response time threshold.

[0141] If the conditions for a dangerous level are not met, but any of the following conditions are met, it is determined to be a warning level: the predicted minimum distance is less than the warning distance threshold and greater than or equal to the dangerous distance threshold; the safe braking margin time is less than the warning time threshold and greater than or equal to the emergency response time threshold.

[0142] If the conditions for neither the danger level nor the warning level are met, it is judged to be at the safe level.

[0143] S104.14, Based on the risk level of all pedestrians, generate the final driving status decision according to the preset decision rules.

[0144] If a pedestrian of high risk is present, an emergency braking decision will be generated; that is, the highest priority is to avoid an immediate collision, the system will trigger maximum deceleration braking and activate relevant warnings.

[0145] If there are no pedestrians at risk level, but there are pedestrians at the warning level, a deceleration decision will be generated. If there are risks that need attention, by actively and smoothly reducing the vehicle speed, the minimum predicted distance and safe braking margin time can be increased, reserving buffer space for possible risk escalation, while improving comfort.

[0146] If all pedestrians are at the safe level, a decision to maintain the current state is generated; if the current environmental risks are controllable, the vehicle can maintain its original speed and planned route.

[0147] S104.2, the cooperating vehicle generates its own driving status decision based on the prediction information of each master vehicle.

[0148] The predicted trajectories of pedestrians under the responsibility of each master vehicle are transformed into the vehicle's coordinate system to obtain the transformed local predicted trajectory. Based on the local predicted minimum distance and local safe braking margin time corresponding to the pedestrian, a risk level is assigned to the pedestrian according to the judgment rules. Based on the risk levels of all pedestrians, the final driving state decision is generated according to the preset decision rules.

[0149] The specific process of this step is the same as the corresponding step in step S104.1, and will not be repeated here.

[0150] Example of collaborative decision-making in a perfect matching scenario: At an unobstructed open intersection, vehicle A (V_A) and vehicle B (V_B) approach from different directions. Pedestrians P_1 and P_2 are both located in the center of the intersection, and both vehicles can clearly and reliably observe these two pedestrians.

[0151] After pedestrian re-identification by the roadside equipment, it is confirmed that both V_A and V_B observe P_1 and P_2, and it is determined to be a "perfectly matched scenario", that is, all pedestrians can be observed by all vehicles. Based on the suitability scores such as distance and viewing angle, it is dynamically decided to assign V_A as the main vehicle for P_1 and V_B as the main vehicle for P_2.

[0152] V_A (as the main vehicle of P_1): Based on multi-view data shared by itself and V_B, trajectory prediction is performed on P_1 to generate prediction information; based solely on the prediction information of P_1 under its responsibility, its risk level (e.g., warning level) is calculated; P_2 prediction information shared by V_B (P_2 master vehicle) is received; its risk level (e.g., safety level) is independently calculated based on the P_2 prediction information; its own decision is generated: the overall decision remains deceleration. Considering the overall risk (warning level + safety level), the decision is generated: deceleration is executed.

[0153] V_B (as the master vehicle of P_2): Similarly, it predicts the trajectory of P_2 based on its own data and the data shared with V_A; it calculates the risk level (e.g., safety level) of P_2 based on the prediction information of P_2 it is responsible for; it receives the prediction information of P_1 shared by V_A (P_1 master vehicle); it independently calculates the risk level (e.g., safety level) of P_1 based on the prediction information of P_1; and generates its own decision: the overall decision is still to decelerate. The overall risk (safety level + safety level) generates the decision: maintain.

[0154] Examples of collaborative decision-making in scenarios with blind spots: At a complex multi-lane intersection, there are four vehicles: vehicle A (V_A), vehicle B (V_B), vehicle C (V_C), and vehicle D (V_D). There are three pedestrians: pedestrian P_X (intending to cross the west side pedestrian crossing from north to south), pedestrian P_Y (intending to cross the north side pedestrian crossing from west to east), and pedestrian P_Z (intending to cross the east side pedestrian crossing from south to north).

[0155] V_A (traveling from north to south): P_X and P_Z can be observed, but P_Y, which is obscured by a large billboard, cannot be observed.

[0156] V_B (traveling from west to east): P_Y can be observed, but P_X, which is obscured by the bus, and P_Z, which is obscured by the bushes, cannot be observed.

[0157] V_C (traveling from south to north): P_Y and P_Z can be observed, but P_X, which is obscured by the V_A vehicle body, cannot be observed.

[0158] V_D (traveling from east to west, in the outermost lane): P_X can be observed, but P_Y and P_Z, which are obscured by the central divider, cannot be observed.

[0159] After pedestrian re-identification by the roadside equipment, V_D, which is closest to P_X and has the most direct view, is assigned as the master vehicle; V_B, with the most stable observations and the highest data quality, is assigned as the master vehicle for P_Y; and V_C is assigned as the master vehicle for P_Z. V_A was not assigned as the master vehicle for any pedestrian in this round of coordination; its role was that of a coordinating vehicle. V_B, V_C, and V_D are all master vehicles.

[0160] The decision-making process of the main vehicle V_B (responsible for P_Y): As the master vehicle of P_Y, V_B acquires its own and V_C's observation data to predict the trajectory of P_Y. V_B receives prediction information shared by other master vehicles via V2X: P_X prediction information shared by V_D (the master vehicle of P_X), and P_Z prediction information shared by V_C (the master vehicle of P_Z). V_B comprehensively processes all pedestrian information: based on its own predicted P_Y, it calculates the risk as a warning level; based on V_D's predicted P_X, after coordinate transformation, it calculates the local risk, finding P_X to be in V_B's blind spot to the side and rear, with no trajectory conflict, thus the risk is at a safe level; based on V_C's predicted P_Z, after coordinate transformation, it calculates the local risk, finding P_Z to be far ahead of V_B, also at a safe level. V_B combines the risks (warning level + safe level + safe level) and generates a decision: to decelerate.

[0161] The decision-making process of the main vehicle V_C (responsible for P_Z): V_C, acting as the master vehicle of P_Z, acquires its own and V_A's observation data to predict the trajectory of P_Z. V_C receives prediction information shared by other master vehicles: V_D (P_X) and V_B (P_Y). V_C performs comprehensive processing: its own predicted risk for P_Z is at the safe level; after conversion and calculation based on the P_X predicted by V_D, it finds that P_X crosses the blind spot to the left front of V_C, indicating a potential trajectory conflict, and the risk is at the warning level; after conversion and calculation based on the P_Y predicted by V_B, the risk is at the safe level. V_C considers the overall risk (safe level + warning level + safe level) and generates a decision: to decelerate.

[0162] The decision-making process of the main vehicle V_D (responsible for P_X): As the master vehicle of P_X, V_D acquires its own and V_A's observation data to predict the trajectory of P_X. V_D receives prediction information shared by other master vehicles: V_B (P_Y) and V_C (P_Z). V_D comprehensively processes the data: its own predicted risk of P_X is classified as hazardous (P_X is crossing the lane where V_D is located); after conversion and calculation based on the P_Y predicted by V_B, the risk is classified as safe; after conversion and calculation based on the P_Z predicted by V_C, the risk is classified as safe. V_D calculates the total risk (hazard level + safe level + safe level) and generates a decision: execute emergency braking.

[0163] The decision-making process of the cooperative vehicle V_A: V_A is not assigned as the primary vehicle; its role is that of a cooperating vehicle. V_A receives prediction information shared by all primary vehicles via V2X: V_D (P_X), V_B (P_Y), and V_C (P_Z). V_A transforms the predicted trajectories of all pedestrians into its own coordinate system and independently calculates the local risk for each pedestrian: For P_X predicted by V_D, which V_A also observes, the risk is calculated as a warning level after verification; for P_Y predicted by V_B, which is a pedestrian in V_A's blind spot, the risk is calculated as a warning level after transformation and calculation, indicating an intersection risk with V_A; for P_Z predicted by V_C, which V_A also observes, the risk is calculated as a safe level. V_A integrates all local risks (warning level + warning level + safe level) and generates a decision: to decelerate.

[0164] Ultimately, V_D (P_X master vehicle) took the most urgent braking measures due to facing the danger directly. V_B (P_Y master vehicle) and V_C (P_Z master vehicle) both took deceleration measures because there was a warning risk of pedestrians or pedestrians in their blind spots. Although V_A (cooperative vehicle) was not a master vehicle, it successfully identified two warning-level risks (including its own blind spot P_Y) based on the complete and high-quality predictive information shared by all master vehicles, and also made a deceleration decision.

[0165] The foregoing has described in detail an embodiment of a vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition. Based on the vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition described in the above embodiment, this invention also provides a vehicle cooperative control system based on vehicle-road cooperative multi-target pedestrian recognition corresponding to the method.

[0166] Figure 2 This is a schematic block diagram of a vehicle cooperative control system based on multi-target pedestrian recognition in vehicle-road cooperative manner, provided as an embodiment of the present invention. Figure 2 As shown, the system includes multiple vehicles and roadside equipment.

[0167] Multiple vehicles: used to collect pedestrian image data within their perception range, extract local perception information from the pedestrian image data, and upload the local perception information to the roadside equipment; the local perception information includes at least pedestrian appearance features and local location information.

[0168] Roadside equipment: used to re-identify pedestrians by determining whether the perception information from different vehicles corresponds to the same actual pedestrian. For each pedestrian that is collected, a primary vehicle is dynamically determined from the vehicles that can observe the pedestrian; vehicles that are not responsible for any pedestrians are recorded as cooperative vehicles.

[0169] Each master vehicle is used to analyze the motion status and predict the behavior of the pedestrians under its responsibility, generating prediction information containing the predicted trajectory of that pedestrian.

[0170] The master vehicle generates its own driving status decision based on its own prediction information and the prediction information of other master vehicles, while the cooperating vehicle generates its own driving status decision based on the prediction information of each master vehicle.

[0171] The vehicle cooperative control system based on vehicle-road cooperative multi-target pedestrian recognition in this embodiment is used to implement the aforementioned vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition. Therefore, the specific implementation of this system can be found in the embodiment section of the vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition mentioned above. Thus, the specific implementation can be referred to the description of the corresponding embodiments, and will not be elaborated here.

[0172] Furthermore, since the vehicle cooperative control system based on vehicle-road cooperative multi-target pedestrian recognition in this embodiment is used to implement the aforementioned vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition, its function corresponds to the function of the above method, and will not be repeated here.

[0173] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A vehicle cooperative control method based on vehicle-road cooperative multi-target pedestrian recognition, characterized in that, Includes the following steps: Multiple vehicles collect pedestrian image data within their respective perception ranges, extract local perception information from the pedestrian image data, and upload the local perception information to the roadside equipment; the local perception information includes at least pedestrian appearance features and local location information; The roadside equipment performs pedestrian re-identification by determining whether the perception information from different vehicles corresponds to the same actual pedestrian; based on the re-identification results, for each pedestrian collected, a master vehicle is dynamically determined from the vehicles that can observe the pedestrian, and vehicles that are not responsible for any pedestrian are recorded as cooperative vehicles. Each master vehicle performs motion state analysis and behavior prediction for the pedestrians it is responsible for, generating prediction information containing the predicted trajectory of that pedestrian; The master vehicle generates its own driving status decision based on its own prediction information and the prediction information of other master vehicles, while the cooperating vehicle generates its own driving status decision based on the prediction information of each master vehicle.

2. The vehicle cooperative control method according to claim 1, characterized in that, Extracting local perception information from pedestrian image data specifically includes: The acquired real-time image frames are processed by an object detection algorithm to locate all pedestrian targets in the image, and a bounding box and corresponding detection confidence score are output for each target. For each detected pedestrian target, the image region within its bounding box is extracted as the region of interest. This region is then input into a pre-trained lightweight convolutional neural network for multi-dimensional feature extraction. The extracted multi-dimensional features include at least clothing category features, color distribution features, and accessory features. Based on the pixel coordinates of the pedestrian target's bounding box in the image and the vehicle's inherent parameters, calculate the pedestrian's local position information relative to the vehicle's coordinate system; this local position information includes at least the pedestrian's relative lateral distance, relative longitudinal distance, and azimuth angle with the vehicle. The extracted multi-dimensional features are concatenated to form a lightweight feature vector. This lightweight feature vector is then associated and encapsulated with local location information, target detection timestamp, and temporary local identifier of the pedestrian target to form a data packet containing the local perception information.

3. The vehicle cooperative control method according to claim 1, characterized in that, Pedestrian re-identification is performed by determining whether perception information from different vehicles corresponds to the same actual pedestrian. Specifically, this includes: The local location information in the data packet is uniformly transformed into the global coordinate system maintained by the roadside equipment to obtain the global location estimate of each observed pedestrian. The lightweight feature vector is decoded to obtain the clothing category feature vector, color distribution feature vector, and accessory feature vector; For all pedestrian observation data received at the current moment, perform a matching operation for any two pedestrian observation data from different vehicles: a) Retrieve three parallel CSPNet residual network branches. The first, second, and third branches process the clothing category feature vector, color distribution feature vector, and accessory feature vector, respectively. During forward propagation of each branch, feature maps from the target intermediate layer of the other two branches are introduced through cross-branch connections. The introduced feature maps are concatenated or element-wise added with the residual edge output of the current branch to output their respective fused feature vectors. Calculate the Euclidean distance between the corresponding fused feature vectors and map the similarity score of the corresponding dimension based on this distance. b) Calculate the physical distance based on the global position estimate after the transformation of the two observation data, and determine whether the physical distance is within the preset movement speed range of the pedestrian based on the timestamp difference between the two data. Map the spatial consistency score according to the judgment result. c) Determine whether two observations correspond to the same actual pedestrian based on spatial consistency score and similarity scores in each dimension.

4. The vehicle cooperative control method according to claim 3, characterized in that, Based on the re-identification results, for each pedestrian collected, a primary vehicle is dynamically determined from the vehicles that can observe that pedestrian. Specifically, this includes: The roadside equipment maintains a list of vehicles that can currently observe each pedestrian for each pedestrian whose data is collected; Based on at least one of the following factors—the relative distance between each candidate vehicle and the pedestrian, the quality of the reported observation data, and the historical observation stability—a suitability score is calculated for each candidate device. The observation data quality is characterized by the target detection confidence and the observation perspective of the candidate device on the pedestrian. The historical observation stability is characterized by the number of consecutive frames or frequency in which the candidate device has successfully reported pedestrian data within a preset time period. The candidate vehicle with the highest suitability score will be designated as the pedestrian's current primary vehicle.

5. The vehicle cooperative control method according to claim 1, characterized in that, Each master vehicle performs motion state analysis and behavior prediction for the pedestrians it is responsible for, generating prediction information containing the predicted trajectory of that pedestrian, specifically including: The main vehicle acquires multi-view observation data of itself and other vehicles on the same pedestrian. Based on the pedestrian's torso target point as the rotation center, the pedestrian image is rotated and transformed to unify the multi-view observation data to a standard analysis perspective. From the unified data, the gait temporal features of the pedestrian are extracted. The gait temporal features include at least the changes in the angle between the leg bone points, the step frequency, and the step length. Based on the extracted gait temporal features, the future predicted position of the pedestrian is calculated using a temporal prediction model, generating a predicted trajectory. The temporal prediction model is expressed as follows: In the formula, For the predicted step size of the next time step, An adaptive weighting coefficient that is positively correlated with the pedestrian's current speed. These are the model coefficients. This represents the change in hip angle at the current moment. This represents the predicted positions of the pedestrian at time t and time t+1 in the global coordinate system. The predicted direction of movement is obtained by a moving average of historical position sequences.

6. The vehicle cooperative control method according to claim 1, characterized in that, The master vehicle generates its own driving status decision based on its own prediction information and the prediction information of other master vehicles, specifically including: Transform the predicted trajectories of pedestrians under the responsibility of other master vehicles into the vehicle's coordinate system to obtain the transformed local predicted trajectories; For each pedestrian collected, the predicted minimum distance and safe braking margin time are calculated based on their predicted trajectory. Based on the predicted minimum distance and safe braking margin time corresponding to the pedestrian, a risk level is assigned to the pedestrian according to the judgment rules; Based on the risk level of all pedestrians, the final driving status decision is generated according to the preset decision rules.

7. The vehicle cooperative control method according to claim 6, characterized in that, For each pedestrian captured, the predicted minimum distance and safe braking margin time are calculated based on their predicted trajectory, specifically including: From the pedestrian's predicted trajectory, select the point that is closest to the vehicle's expected path space within a fixed future time window as the nearest interaction point, and record the predicted location and arrival time of that point. Based on the vehicle's current position, current speed, and preset deceleration capability, calculate the predicted minimum distance and safe braking margin time, including: a) The predicted minimum distance is the Euclidean distance between the vehicle's current position and the predicted position when the vehicle travels at a constant speed to the position at the arrival time. b) Calculate the vehicle's braking distance based on the vehicle's current speed and maximum safe deceleration. Calculate the time required for the vehicle to come to a complete stop from the start of braking, and the distance the pedestrian moves during this time. Starting from the vehicle's current position, extend backward along the current direction of travel. The spatial area defined by the sum of the vehicle's braking distance, the distance the pedestrian moves during this time, and the preset static safety margin constitutes a safety envelope. Predict the time required for the pedestrian's trajectory to first enter the safety envelope from the current position, moving at its current speed and direction. This time is the safe braking margin time.

8. The vehicle cooperative control method according to claim 6, characterized in that, Risk levels include hazardous, warning, and safe levels; Based on the risk level of all responsible pedestrians, the final driving status decision is generated according to preset decision rules, including: if there are dangerous pedestrians, an emergency braking decision is generated; if there are no dangerous pedestrians but there are warning pedestrians, a deceleration decision is generated; if all pedestrians are at the safe level, a decision to maintain the current state is generated.

9. The vehicle cooperative control method according to claim 1, characterized in that, The collaborative vehicle generates its own driving status decision based on the prediction information from each master vehicle, specifically including: The predicted trajectories of pedestrians under the responsibility of each master vehicle are transformed into the vehicle's coordinate system to obtain the transformed local predicted trajectories. Based on the local predicted minimum distance and local safe braking margin time corresponding to the pedestrian, a risk level is assigned to the pedestrian according to the judgment rules. Based on the risk level of all pedestrians, the final driving status decision is generated according to the preset decision rules.

10. A vehicle cooperative control system based on vehicle-road cooperative multi-target pedestrian recognition, characterized in that, include: Multiple vehicles: used to collect pedestrian image data within their perception range, extract local perception information from the pedestrian image data, and upload the local perception information to the roadside equipment; the local perception information includes at least pedestrian appearance features and local location information; Roadside equipment: used to re-identify pedestrians by determining whether the perception information from different vehicles corresponds to the same actual pedestrian. For each pedestrian that is collected, a primary vehicle is dynamically determined from the vehicles that can observe the pedestrian; vehicles that are not responsible for any pedestrians are recorded as cooperative vehicles. Each master vehicle is used to analyze the motion status and predict the behavior of the pedestrians under its responsibility, and generate prediction information containing the predicted trajectory of the pedestrian. The master vehicle generates its own driving status decision based on its own prediction information and the prediction information of other master vehicles, while the cooperating vehicle generates its own driving status decision based on the prediction information of each master vehicle.