Method and apparatus for identifying a type of vehicle stop
By acquiring traffic scene information and traffic light timing features, and using a recognition model to identify vehicle parking types, the problem of insufficient recognition accuracy in existing technologies is solved, and the ability to identify long-term parked vehicles is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING VOYAGER TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies, especially for identifying long-term parked vehicles, suffer from insufficient learning ability regarding interactions between traffic participants and road conditions, resulting in low accuracy.
By acquiring environmental information of the traffic scene, the coding unit in the recognition model is used to determine the object characteristics of traffic participants, and combined with the timing characteristics of traffic lights, the decoding unit is used to determine the vehicle's stopping type, especially the expected stopping time.
It improves the accuracy of vehicle parking type identification, especially the ability to identify long-term parked vehicles, and enhances the interactivity and robustness of the model.
Smart Images

Figure CN122116649A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for identifying vehicle parking types. Background Technology
[0002] With the rapid development of transportation systems, vehicle recognition and motion detection are being widely applied in areas such as road traffic safety and autonomous driving. For example, in transportation systems, it is necessary to detect and track moving vehicles for applications such as traffic flow statistics. Correspondingly, determining whether a target vehicle intends to leave in the future based on traffic participant and road information can play an important role in informing the lane-changing intentions of moving vehicles in their decision-making path planning. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for identifying vehicle parking types is provided. The method includes: acquiring environmental information associated with a traffic scene; using an encoding unit in an identification model, determining object features associated with at least one traffic participant in the traffic scene based on the environmental information, the at least one traffic participant including a target vehicle; and using a decoding unit in the identification model, determining the parking type of the target vehicle based on the object features and temporal features associated with traffic lights in the traffic scene, the parking type being associated with the expected parking time of the target vehicle.
[0004] In a second aspect of this disclosure, an apparatus for identifying vehicle parking types is provided. The apparatus includes: an environmental information acquisition module configured to acquire environmental information associated with a traffic scene; a feature determination module configured to use an encoding unit in an identification model to determine object features associated with at least one traffic participant in the traffic scene based on the environmental information, the at least one traffic participant including a target vehicle; and a parking type determination module configured to use a decoding unit in the identification model to determine the parking type of the target vehicle based on the object features and temporal features associated with traffic lights in the traffic scene, the parking type being associated with the expected parking time of the target vehicle.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.
[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 2 A schematic diagram of an example architecture for training a machine learning model is shown according to some embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram of an example architecture of an applied trained machine learning model according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic diagram of an example process for a travel service according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A schematic structural block diagram of an example device for a travel service according to certain embodiments of the present disclosure is shown; and
[0015] Figure 6 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0017] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0019] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0020] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0021] As briefly mentioned earlier, determining whether a target vehicle intends to leave based on traffic participant and road information can serve as an important reference for lane-changing intentions in the decision-making and path planning of moving vehicles. Currently, most long-stop recognition systems are based on ensemble learning algorithms. For example, algorithms such as extreme gradient boosting and ensemble learning algorithms can be used to determine whether a vehicle will remain stationary within a predetermined time period. However, this approach has limited learning capabilities regarding interactions between traffic participants and between traffic participants and road conditions (e.g., map information).
[0022] Embodiments of this disclosure propose an improved scheme for identifying vehicle stops. According to various embodiments of this disclosure, environmental information associated with a traffic scene is acquired. Further, using an encoding unit in the identification model, object features associated with at least one traffic participant in the traffic scene are determined based on the environmental information; the at least one traffic participant includes a target vehicle. Then, by utilizing a decoding unit in the identification model, the stop type of the target vehicle is determined based on the object features and temporal features associated with traffic lights in the traffic scene; the stop type is associated with the expected stop time of the target vehicle.
[0023] In this way, embodiments of the present disclosure can improve the accuracy of identifying the parking type of a vehicle (e.g., whether it is a long-term parking vehicle or a short-term parking vehicle).
[0024] Example Environment
[0025] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. Figure 1 In environment 100, it is desirable to train and use a recognition model 130, which is configured for various application environments. For example, when the recognition model 130 is used to identify whether a vehicle is parked for a long time (i.e., whether the vehicle stays within a predetermined time period), it can output a result on whether the vehicle is parked for a long time based on some feature information of the vehicle input by the user.
[0026] like Figure 1 As shown, environment 100 includes model training system 150. Figure 1 The upper part illustrates the model training phase, and the lower part illustrates the model application phase. Before training, the parameter values of the recognition model 130 can have initial values or pre-trained parameter values obtained through a pre-training process. The recognition model 130 can be trained via vectorized encoding, and its parameter values can be updated and adjusted during training. After the recognition model 130 is trained, recognition model 130′ is obtained. At this point, the parameter values of recognition model 130′ have been updated, and based on the updated parameter values, recognition model 130 can be used in the model application phase to identify whether a vehicle is parked for an extended period.
[0027] During the model training phase, the recognition model 130 can be trained using a model training system 150 based on a sample feature set 110 that includes multiple sample features 112 used to train the recognition model 130. Here, sample features 112 may include first sample features 120, second sample features 122, and / or third sample features 123. For example, for a task indicating whether a vehicle is a long-term parked vehicle, sample features 112 may include first sample features 120 indicating a vehicle and its surrounding traffic participants. Sample features 112 may also include second sample features 122 indicating map information corresponding to a vehicle's surroundings. Additionally, sample features 112 may also include third sample features 123 indicating traffic lights linked to map information. First sample features 120, second sample features 122, and / or third sample features 123 can be used to train the recognition model 130. Specifically, the training process can be performed iteratively using a large number of training samples. After training is complete, the recognition model 130 may include knowledge about the task to be processed. During the model application phase, the recognition model 130′ (which at this time has trained parameter values) can be used to perform the corresponding task. For example, it can receive multiple features 142 indicating whether a vehicle is a long-term parked vehicle in the task of identifying whether the vehicle is a long-term parked vehicle, and output the parking type 144 for identifying whether the vehicle is a long-term parked vehicle.
[0028] exist Figure 1 In this context, the model training system 150 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0029] It should be understood that Figure 1 The components and arrangements shown in environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. Implementations of this disclosure are not limited in this respect.
[0030] To facilitate understanding, the following description first refers to an example to illustrate the application scenario of this disclosure, but this is merely exemplary and not intended to limit the scope of the disclosure. For ease of description, this example will be presented from the perspective of any vehicle currently in motion. Vehicle A, currently in motion, can obtain current road conditions through its corresponding server-side device. For example, vehicle A can decide whether to change lanes based on the information about upstream vehicles provided by the server-side device. Therefore, in this embodiment of the disclosure, the server-side device can classify the target vehicle into two categories based on whether it is a long-term parked vehicle (e.g., vehicles parked on the roadside, vehicles involved in traffic accidents, etc., can all be considered long-term parked vehicles), providing a reference for downstream vehicles (e.g., vehicle A) when changing lanes.
[0031] The following description, with reference to the accompanying drawings, illustrates how a server device performs binary classification to determine whether a target vehicle is a long-term parked vehicle. In some embodiments, the server device may utilize a trained recognition model to determine whether a target vehicle is a long-term parked vehicle (i.e., whether the target vehicle remains parked within a predetermined time period). It should be understood that a model training system 150 and a model application system 160 may be deployed in the server device. Hereinafter, exemplary embodiments will be described primarily with respect to the server device. The exemplary embodiments are described using model training system 150 and model application system 160. It should be understood that the actions described with respect to the server device may be performed by model training system 150 on the server device, or by model training system 150 in conjunction with its server (e.g., a server). Correspondingly, the actions described with respect to the server device may be performed by model application system 160 on the server device, or by model application system 160 in conjunction with its server (e.g., a server).
[0032] Example training process
[0033] The following will first refer to Figure 1 and Figure 2 This describes the training process of the recognition model. Figure 2 A schematic diagram of an example architecture 200 for training a recognition model according to some embodiments of the present disclosure is shown. Hereinafter, the example embodiments will be described primarily with respect to a model training system 150 on a server device. It should be understood that the actions described with respect to the model training system 150 can be performed by a recognition model 130 on the model training system 150, or can be performed by the recognition model 130 in conjunction with its server (e.g., a server).
[0034] In embodiments of this disclosure, the model training system 150 acquires sample data for training a recognition model. The model training system 150 processes the sample data using encoding units to generate training global features.
[0035] like Figure 2As shown, the sample data may include sample feature 120, sample feature 122, and sample feature 123. Sample feature 120 can be used to indicate the characteristics of the training vehicle and its surrounding traffic participants, and can be represented as x. agent In some examples, sample feature 120 may include historical trajectory information of the traffic participant, which may be represented in coordinate form. Sample feature 120 may include attribute information of the traffic participant, such as whether the traffic participant can be a vehicle, pedestrian, etc. In some examples, if the traffic participant is a vehicle, it is further determined that the traffic participant can be a different type of vehicle (e.g., bus, sedan, truck, etc.). In some examples, sample feature 120 may include the status of the traffic participant's lights, such as the taillights of the traffic participant. In some examples, sample feature 120 may also include 3D bounding box parameters for the traffic participant, which can be used to indicate the volume size of the traffic participant. In some examples, sample feature 120 may also include motion parameters for the traffic participant, such as the traffic participant's orientation, heading, speed, etc.
[0036] In some embodiments, sample feature 120(x) agent This can be input into the recognition model 130 as a tensor matrix. In some examples, x agent The input to the recognition model 130 can be a tensor matrix of (Batch_size, Num_agents, History_length, Feature_size). Batch_size indicates the batch size of the deep learning network, Num_agents indicates the number of nearby traffic participants, History_length indicates the length of the historical frames input, and Feature_size indicates the number of features. In some examples, tensors include vectors (1-dimensional), matrices (2-dimensional), etc., and are a collective term for arrays of different dimensions.
[0037] In some embodiments, sample feature 122 can be used to indicate features corresponding to map information around the training vehicle, and can be represented as x. map In some examples, sample feature 122(x) map This can include features of map location information, features of attribute information (e.g., features of lane type information), features of information associated with traffic lights, etc. Correspondingly, x mapThe input to the recognition model 130 can be a tensor matrix of (Batch_size, Num_map_objects, Segment_num, Feature_size). Num_map_objects represents the number of map elements around the target vehicle, and Segment_num represents the number of vectors that each map element is segmented into.
[0038] In some embodiments, sample feature 123 can be used to indicate temporal features associated with traffic lights in a traffic scene, and can be represented as x tl In some examples, x tl The input to the recognition model 130 can be a tensor matrix of (Batch_size, Num_map_objects, History_length, Feature_size).
[0039] Continue to refer to Figure 2 The example architecture 200 shown has a model training system 150 based on sample features (x). agent 120. Sample characteristics (x) map 122. Sample characteristics (x) tl )123, by utilizing the local encoder 210, the local features (PF) corresponding to sample features 120, 122, and 123 are obtained respectively. agent 212. Local Features (PF) map 213. Local Features (PF) tl )214.
[0040] That is, the model training system 150 trains the model based on the sample features (x) agent )120, Obtain Local Features (PF) agent )212 can be represented as: PF agent =Encoder agent (x agent The model training system 150 is based on sample features (x) map )122, Obtain Local Features (PF) map )213 can be represented as: PF map =Encoder map (x map Accordingly, the model training system 150 trains the model based on the sample features (x). tl )123, Obtain Local Features (PF) tl )214 can be represented as: PF tl =Encoder tl (x tl ).
[0041] Subsequently, the model system 150 can convert local features (PF) agent The target part in )212, and local features (PF) map The target portion in )213 is replaced with preset values to determine the mask features. That is, the model training system 150 can use the obtained local features (PF) agent 212. Local Features (PF) map )213, respectively set to zero according to a predetermined ratio (e.g., XX%), and then determine the mask features (e.g., mask feature 215, mask feature 216, and mask feature 217).
[0042] Correspondingly, the model training system 150 can also train local features (PF). agent )212 and local features (PF) map )213 After being zeroed out according to a predetermined ratio, a normalized location code (scaling identifier, SI) is spliced in. Understandably, the model application system 160 performs normalization processing on the location information of at least one traffic participant based on the historical location distribution of a set of traffic participants in the traffic scenario to determine the location feature 218 (SI).
[0043] In some examples, SI can indicate the relative position of any location. In some embodiments, for local features (PF) agent The normalized encoding concatenated by )212 can be represented as SI agent Understandably, the model training system 150 can take the trajectory of the last frame corresponding to the traffic participant and normalize it according to the trajectories of all vehicles in the surrounding environment. For example, based on the distribution of vehicles in the traffic scene, the absolute position of vehicle A can be normalized to a predetermined interval (e.g., interval (0,1), interval (1,-1)), thereby obtaining the relative position of vehicle A.
[0044] In some embodiments, for local features (PF) map The normalized encoding concatenated by )213 can be represented as SI map Understandably, the model training system 150 can take the center point of map elements and normalize them according to all map elements in the surrounding environment.
[0045] In some embodiments, the model training system 150 can be based on local features (PF) that have been zeroed out according to a predetermined ratio. agent 212. Local Features (PF) map 213. Local Features (PF) tl )214, and the normalized SI agent SI mapBy utilizing the global feature encoder 220, training global features (GF) are obtained. global )221. This can be represented as: GF global =Encoder global (Mask(PF agent )+S.Iagent,Mask(PF map )+SI map ,PF tl ).
[0046] In some embodiments, the model training system 150 may utilize a first decoding unit to process training global features to determine the predicted docking type. Then, the model training system 150 determines a training loss based on a first difference between the predicted docking type and a reference docking type.
[0047] like Figure 2 As shown, the model training system 150 can use the first decoder 223 included in the first decoding unit to process the training global features 221 to determine the predicted parking type 224. That is, based on the training global features 221, the model training system 150 can determine a binary classification of whether the training vehicle and other vehicles are long-term parking vehicles by using the first decoder 223. This binary classification can be represented, for example, as: Decoder parked_car (GF agent_vehicle ).
[0048] Then, the model training system 150 can be based on the predicted parking type 224 (Decoder) corresponding to the training vehicle. parked_car (GF agent_vehicle The training vehicle's corresponding actual parking type (Label) parked_car The difference between ) is used to determine the training loss (L) parked car ), and then train the recognition model 130 based on this loss. This can be expressed as:
[0049] L parked car =
[0050] CrossEntropyLoss(Decoder parked_car (GF agent_vehicle ),Label parked_car ).
[0051] In some embodiments, the model training system 150 may also utilize a second decoding unit to process training global features to determine the predicted trajectory associated with at least one participant. Furthermore, the model training system 150 may train the recognition model 130 based on a second difference between the predicted trajectory and a reference trajectory associated with at least one participant.
[0052] like Figure 2 As shown, the model training system 150, based on training global features 221, can determine the predicted trajectory 226 associated with traffic participants by utilizing the second decoder 225 in the second decoding unit. The predicted trajectory 226 can be represented as: Decoder trajectory_predict (GF agent_vehicle ).
[0053] Then, the model training system 150 can train the model based on the predicted trajectory 226. and the real trajectory label associated with traffic participants trajectory The difference between them is used to determine the training loss. Then, the recognition model 130 is trained based on this loss. This loss can be expressed as:
[0054] In some embodiments, the second decoding unit is configured to acquire mask features. In some embodiments, the mask features are determined by replacing the target portion of trained global features with preset values. Figure 2 As shown, the second decoding unit includes a third decoder 227, which is configured to acquire mask features 215, mask features 216, and / or mask features 217.
[0055] In some embodiments, the model training system 150 generates predicted features for the target portion based on mask features. Further, the model training system 150 can train the recognition model 130 based on a third difference between the predicted features and the target features corresponding to the training global features and the target portion.
[0056] like Figure 2 As shown, the model training system 150 can determine the prediction feature 228 corresponding to the mask feature 215 by using the third decoder 227. The prediction feature 228 can be represented as: (GF mask ).
[0057] Then, the model training system 150 can be based on the predicted features 228 (Decoder) graph_completion (GF mask )) and the target features (PF) corresponding to the target part mask The difference between ) is used to determine the training loss (L) graph_completion ), and then train the recognition model 130 based on this loss. As an example, the training loss can be expressed as: L graph_completion =L1Loss(Decoder) graph_completion (GF mask ),PF mask ).
[0058] In some embodiments, the model training system 150 may be based on the training loss (L parked car Training loss and training loss (L graph_completion The weights are then applied to train the recognition model 130. This can be represented as: Where α is the loss function L parked car The weights, where β is the loss function. The weights, where γ is the loss function. The weights α, β, and γ are updated during the training of the recognition model 130. When the weights α, β, and γ are the target values, the prediction accuracy of the recognition model 130 is higher. In some examples, if α·L parked car For γ·L graph_completion Ten times, for If the weight is five times that of the model, the accuracy of the prediction results of model 130 will be higher. It should be understood that the specific weight values mentioned above are only illustrative.
[0059] Based on the training process described above, the embodiments of this disclosure can improve the recognition ability, interactivity, and robustness of the recognition model by collecting diverse sample data. Furthermore, by training the recognition model based on temporal features associated with traffic lights, the model's ability to recognize traffic lights can be simultaneously enhanced.
[0060] Example application process
[0061] The following will be further referenced Figure 3 This section describes a schematic diagram of an example architecture 300 for applying a trained machine learning model according to some embodiments of the present disclosure. Hereinafter, the example embodiments will be described primarily with respect to a model application system 160 on a server device. It should be understood that the actions described with respect to the model application system 160 may be performed by a recognition model 130' on the model application system 160, or may be performed by the recognition model 130' in conjunction with its server (e.g., a server).
[0062] In embodiments of this disclosure, the model application system 160 on the server device acquires environmental information related to the traffic scene. In some examples, the environmental information may include relevant information about the target vehicle and surrounding traffic participants, map information corresponding to the area around the target vehicle, and / or timing information of traffic lights bound to the map information.
[0063] In some examples, information about the target vehicle and surrounding traffic participants may include the traffic participants' historical trajectory information, which may be represented in coordinate form. This information may also include the traffic participants' attribute information, such as whether they are vehicles, pedestrians, etc. In some examples, if a traffic participant is a vehicle, it may be further determined that the participant can be of different vehicle types (e.g., bus, sedan, truck, etc.).
[0064] In some examples, information about the target vehicle and surrounding traffic participants may include the status of the traffic participants' lights, such as their taillights. In some examples, this information may also include 3D bounding box parameters for the traffic participants, which can be used to indicate their volume. In some examples, this information may also include motion parameters for the traffic participants, such as their orientation, heading, and speed.
[0065] In some embodiments, the map information corresponding to the target vehicle may include map location information, attribute information, information associated with traffic lights, etc. In some examples, the attribute information included in the map information corresponding to the target vehicle may indicate the type of lane. For example, a bus lane, a left-to-right lane, a right-to-left lane, etc. In some examples, the information associated with traffic lights included in the map information corresponding to the target vehicle may indicate the light status information of the traffic lights bound to the map information, etc. In some embodiments, the timing information of the traffic lights bound to the map information may indicate the light status information of the traffic lights bound to key points in the map elements.
[0066] In embodiments of this disclosure, the model application system 160 utilizes encoding units in a trained recognition model and determines object features associated with at least one traffic participant in a traffic scene based on environmental information. In some embodiments, the at least one traffic participant includes a target vehicle.
[0067] In some embodiments, the object characteristics determined by the model application system 160 based on environmental information using the coding unit may include the historical trajectories of traffic participants. For example... Figure 3 The example architecture shown is 300, and the object feature is 311 (x). agent This includes features corresponding to the historical trajectory information of traffic participants. The object features determined by the model application system 160 based on environmental information using the encoding unit may include the motion parameters of the traffic participants. The object features determined by the model application system 160 may also include the vehicle light status of the traffic participants. In some examples, object feature 311(x)agent This includes features corresponding to the vehicle light status information of traffic participants.
[0068] In some embodiments, the object features determined by the model application system 160 based on environmental information using the coding unit may include the location of traffic participants. In some examples, object feature 311(x agent This includes features corresponding to the 3D detection bounding box parameters for traffic participants. Additionally, the object features determined by the encoding unit based on environmental information in the model application system 160 may also include the type of traffic participant. In some examples, object feature 311 (x agent This includes the characteristics corresponding to the type information of traffic participants.
[0069] In some embodiments, the model application system 160 can determine local features corresponding to at least one traffic participant based on environmental information. Then, the model application system 160 performs normalization processing on the location information of at least one traffic participant based on the historical location distribution of a set of traffic participants in the traffic scene to determine location features.
[0070] like Figure 3 As shown, after obtaining object features 311 based on environmental information, the model application system 160 can also use the local encoder 310 to obtain local features 314 (PF) of the object features 311. agent It can be represented as: PF agent =Bncoder agent (x agent Accordingly, the model application system 160 performs normalization processing on the location information corresponding to object feature 311 based on the historical location distribution of a set of traffic participants in the traffic scenario, in order to determine the location feature 319 (which can be represented as SI) for object feature 311. agent Understandably, during the model application phase, after acquiring the local features of object feature 311, a masking operation is not performed to use all the local features to determine the parking type of the target vehicle. This approach avoids information loss.
[0071] In some embodiments, such as Figure 3 As shown, the model application system 160 can obtain time-series features 313(x) based on environmental information. tl It is understandable that temporal features associated with traffic lights can indicate features corresponding to the light status information of traffic lights bound to key points in map elements. In some embodiments, the model application system 160 can utilize the local encoder 310 to obtain local features 316 (PF) of the temporal features 313. tl It can be represented as: PF tl =Encodertl (x tl ).
[0072] In some embodiments, the model application system 160 can also acquire map information associated with the traffic scene. Furthermore, based on the map information, the model application system 160 can determine map features associated with the target vehicle. For example... Figure 3 As shown, the model application system 160 obtains map features 312 (x) based on map information. map After that, the local encoder 310 can be used to obtain the local features 315 (PF) of the map features 312. map It can be represented as: PF map =Encoder map (x map ).
[0073] Additionally, the model application system 160 can also perform normalization processing on the location information corresponding to map feature 312 based on the historical location distribution of a set of traffic participants in the traffic scenario, in order to determine the location feature 319 (which can be represented as SI) for map feature 312. map Understandably, during the model application phase, after acquiring the local features of map feature 312, a masking operation is not performed to use all the local features to determine the parking type of the target vehicle. This approach avoids information loss.
[0074] In some embodiments, the model application system 160 determines global features based on object features, map features, and temporal features. For example... Figure 3 As shown, the model application system 160 can be based on local features 314 (PF) agent ), for object feature 311, position feature 319 (SI) agent ), Local feature 316 (PF) tl ), Local feature 315 (PF) map ), and location feature 319 (SI) for map feature 312. map Global features 317 (GF) are obtained using the global encoder 320. global As an example, global feature 317 can be represented as: GF global =Encoder global (Mask(PF agent )+SI agent Mask(PF) map )+SI map ,PF tl ).
[0075] In some embodiments, the model application system 160 uses a decoding unit in the recognition model to determine the parking type of a target vehicle based on object features and temporal features associated with traffic lights in a traffic scene. In some embodiments, the parking type is associated with the expected parking time of the target vehicle.
[0076] In some examples, the model application system 160 can determine the parking type 144 of the target vehicle based on object feature 311 and temporal feature 313. In other examples, the model application system 160 can determine the parking type 144 of the target vehicle based on object feature 311, location features related to object feature 311, and temporal feature 313.
[0077] In other embodiments, the model application system 160 determines the parking type of the target vehicle based on object features, map features, and temporal features. In some examples, the model application system 160 can determine the parking type 144 of the target vehicle based on object feature 311, map feature 312, and temporal feature 313. In other examples, the model application system 160 can determine the parking type 144 of the target vehicle based on object feature 311, location features for object feature 311, map feature 312, location features for map feature 312, and temporal feature 313.
[0078] In some embodiments, the model application system 160 utilizes a decoding unit to process global features to determine the parking type of the target vehicle. For example... Figure 3 As shown, the model application system 160, based on global features 317, can determine the parking type 144 of the target vehicle (e.g., whether the target vehicle is a long-term parking vehicle) by using the first decoder 223 included in the decoding unit of the recognition model 130′.
[0079] Based on the application process of the recognition model 130′ described above, the embodiments of this disclosure can improve the accuracy of identifying the parking type of a vehicle.
[0080] Example process
[0081] Figure 4 A flowchart of an example process 400 for identifying vehicle parking type according to some embodiments of the present disclosure is shown. Process 400 can be implemented at model application system 160. References are as follows. Figure 1 Describe the process 400.
[0082] like Figure 4 As shown in box 410, the model application system 160 acquires environmental information associated with the traffic scene.
[0083] In box 420, model application system 160 uses coding units in the recognition model to determine object features associated with at least one traffic participant in the traffic scene based on environmental information, the at least one traffic participant including the target vehicle.
[0084] In box 430, the model application system 160 uses the decoding unit in the recognition model to determine the parking type of the target vehicle based on object features and temporal features associated with traffic lights in the traffic scene. The parking type is associated with the expected parking time of the target vehicle.
[0085] In some embodiments, determining object features associated with at least one traffic participant in a traffic scenario based on environmental information includes: determining local features corresponding to at least one traffic participant based on environmental information; performing normalization processing on the location information of at least one traffic participant based on the historical location distribution of a group of traffic participants in the traffic scenario to determine location features; and determining object features associated with at least one traffic participant based on local features and location features.
[0086] In some embodiments, determining the parking type of a target vehicle based on object features and temporal features associated with traffic lights in a traffic scene includes: acquiring map information associated with the traffic scene; generating map features associated with the target vehicle based on the map information; and determining the parking type of the target vehicle based on the object features, map features, and temporal features.
[0087] In some embodiments, determining the parking type of a target vehicle based on object features, map features, and temporal features includes: determining global features based on object features, map features, and temporal features; and processing the global features using a decoding unit to determine the parking type of the target vehicle.
[0088] In some embodiments, the decoding unit is a first decoding unit, and the recognition model is trained based on the following process: processing sample data using an encoding unit to generate training global features; processing the training global features using the first decoding unit to determine the predicted docking type; determining a training loss based on a first difference between the predicted docking type and a reference docking type, wherein the training loss is also based on the prediction results of the training global features by a second decoding unit; and training the recognition model based on the training loss.
[0089] In some embodiments, the prediction result of the second decoding unit includes a predicted trajectory associated with at least one participant, and the training loss is also based on a second difference between the predicted trajectory and the reference trajectory associated with at least one participant.
[0090] In some embodiments, the second decoding unit is configured to: acquire mask features, which are determined by replacing the target portion of the training global features with preset values; and generate a prediction result based on the mask features, wherein the prediction result includes prediction features for the target portion, and the training loss is also based on a third difference between the prediction features and the target features corresponding to the training global features and the target portion.
[0091] In some embodiments, the second decoding unit is disabled during the recommendation phase of the recognition model.
[0092] In some embodiments, the object features include at least one of the following: the traffic participant's historical trajectory; the traffic participant's motion parameters; the traffic participant's headlight status; the traffic participant's location; and / or the type of the traffic participant.
[0093] Example devices and equipment
[0094] Figure 5 A schematic structural block diagram of a device 500 for identifying vehicle parking type according to certain embodiments of the present disclosure is shown. The device 500 may be implemented as or included in the model application system 160. The various modules / components in the device 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0095] As shown in the figure, the device 500 includes an environmental information acquisition module 510, configured to acquire environmental information associated with a traffic scene; a feature determination module 520, configured to use an encoding unit in a recognition model to determine object features associated with at least one traffic participant in the traffic scene based on the environmental information, wherein the at least one traffic participant includes a target vehicle; and a stop type determination module 530, configured to use a decoding unit in a recognition model to determine the stop type of the target vehicle based on the object features and temporal features associated with traffic lights in the traffic scene, wherein the stop type is associated with the expected stop time of the target vehicle.
[0096] In some embodiments, the feature determination module 520 is further configured to: determine local features corresponding to at least one traffic participant based on environmental information; perform normalization processing on the location information of at least one traffic participant based on the historical location distribution of a group of traffic participants in a traffic scene to determine location features; and determine object features associated with at least one traffic participant based on local features and location features.
[0097] In some embodiments, the parking type determination module 530 is further configured to acquire map information associated with a traffic scenario; generate map features associated with a target vehicle based on the map information; and determine the parking type of the target vehicle based on object features, map features, and temporal features.
[0098] In some embodiments, the parking type determination module 530 is further configured to determine global features based on object features, map features, and temporal features; and to process the global features using a decoding unit to determine the parking type of the target vehicle.
[0099] In some embodiments, the decoding unit is a first decoding unit, and the apparatus 500 includes a model training module configured to process sample data using an encoding unit to generate training global features; process the training global features using the first decoding unit to determine a predicted docking type; determine a training loss based on a first difference between the predicted docking type and a reference docking type, wherein the training loss is also based on the prediction result of the training global features by the second decoding unit; and train a recognition model based on the training loss.
[0100] In some embodiments, the prediction result of the second decoding unit includes a predicted trajectory associated with at least one participant, and the training loss is also based on a second difference between the predicted trajectory and the reference trajectory associated with at least one participant.
[0101] In some embodiments, the second decoding unit is configured to: acquire mask features, which are determined by replacing the target portion of the training global features with preset values; and generate a prediction result based on the mask features, wherein the prediction result includes prediction features for the target portion, and the training loss is also based on a third difference between the prediction features and the target features corresponding to the training global features and the target portion.
[0102] In some embodiments, the second decoding unit is disabled during the recommendation phase of the recognition model.
[0103] In some embodiments, the object features include at least one of the following: the traffic participant's historical trajectory; the traffic participant's motion parameters; the traffic participant's headlight status; the traffic participant's location; and / or the type of the traffic participant.
[0104] Figure 6 A block diagram is shown illustrating a computing device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 6 The computing device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 6 The computing device 600 shown can be used to implement Figure 1 The model application system 160 and / or the model training system 150.
[0105] like Figure 6As shown, computing device 600 is in the form of a general-purpose computing device. Components of computing device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage devices 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 600.
[0106] Computing device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to computing device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 600.
[0107] The computing device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0108] The communication unit 640 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 600 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0109] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 600 can also communicate as needed with one or more external devices (not shown) via communication unit 640. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 600, or with any device (e.g., network card, modem, etc.) that enables computing device 600 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0110] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0111] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0112] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0113] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0115] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for identifying vehicle parking type, comprising: Obtain environmental information related to traffic scenarios; Using the coding unit in the recognition model, based on the environmental information, object features associated with at least one traffic participant in the traffic scenario are determined, wherein the at least one traffic participant includes the target vehicle; as well as Using the decoding unit in the recognition model, the parking type of the target vehicle is determined based on the object features and the temporal features associated with the traffic lights in the traffic scene. The parking type is associated with the expected parking time of the target vehicle.
2. The method according to claim 1, wherein determining the object features associated with at least one traffic participant in the traffic scenario based on the environmental information comprises: Based on the environmental information, determine the local features corresponding to the at least one traffic participant; Based on the historical location distribution of a group of traffic participants in the traffic scenario, the location information of at least one traffic participant is normalized to determine location features. as well as Based on the local features and the location features, the object features associated with the at least one traffic participant are determined.
3. The method according to claim 1, wherein determining the parking type of the target vehicle based on the object characteristics and temporal characteristics associated with traffic lights in the traffic scenario includes: Obtain map information associated with the traffic scenario; Based on the map information, map features associated with the target vehicle are generated; as well as The parking type of the target vehicle is determined based on the object features, the map features, and the time series features.
4. The method according to claim 3, wherein determining the parking type of the target vehicle based on the object features, the map features, and the temporal features includes: Global features are determined based on the object features, the map features, and the temporal features; as well as The global features are processed using the decoding unit to determine the parking type of the target vehicle.
5. The method according to claim 1, wherein the decoding unit is a first decoding unit, and the recognition model is trained based on the following process: The encoding unit is used to process sample data to generate training global features; The first decoding unit processes the trained global features to determine the predicted docking type; The training loss is determined based on the first difference between the predicted docking type and the reference docking type, wherein the training loss is also based on the prediction result of the training global feature by the second decoding unit; as well as The recognition model is trained based on the training loss.
6. The method of claim 5, wherein the prediction result of the second decoding unit includes a predicted trajectory associated with the at least one participant, and the training loss is further based on a second difference between the predicted trajectory and a reference trajectory associated with the at least one participant.
7. The method of claim 5, wherein the second decoding unit is configured to: Obtain mask features, which are determined by replacing the target portion of the trained global features with preset values; and The prediction result is generated based on the mask features, wherein the prediction result includes prediction features for the target part, and the training loss is also based on the prediction features and a third difference between the training global features and the target features corresponding to the target part.
8. The method of claim 5, wherein the second decoding unit is disabled during the recommendation phase of the recognition model.
9. The method of claim 1, wherein the object features include at least one of the following: The historical trajectory of traffic participants; Motion parameters of traffic participants; The status of vehicle lights of traffic participants; Location of traffic participants; and / or Types of traffic participants.
10. A device for identifying vehicle parking type, comprising: The environmental information acquisition module is configured to acquire environmental information related to the traffic scenario; The feature determination module is configured to use the coding unit in the recognition model to determine object features associated with at least one traffic participant in the traffic scenario based on the environmental information, wherein the at least one traffic participant includes the target vehicle. as well as The parking type determination module is configured to use the decoding unit in the recognition model to determine the parking type of the target vehicle based on the object features and the temporal features associated with the traffic lights in the traffic scene, wherein the parking type is associated with the expected parking time of the target vehicle.
11. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.