Parking control methods, storage media and vehicles
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]有鉴于此,本公开的目的在于提出一种停车控制方法、存储介质及车辆,以解决现有技术中不能全面理解车辆周围环境导致停车区域识别不准确的问题
[0008]As described above, the parking control method, storage medium, and vehicle disclosed herein acquire environmental data surrounding the vehicle and the vehicle's own state data, and construct a semantic graph based on the environmental and state data. Utilizing a pre-trained cross-modal attention network, the semantic features of each physical region are obtained by fusing the node features corresponding to multiple nodes in the semantic graph, effectively integrating multi-source environmental and state data and enhancing the understanding of complex environments. Using a pre-trained spatiotemporal graph generation model, the temporal features of each physical region are determined based on the semantic features, capturing the dynamic changes of physical regions at different times and providing a temporal dimension basis for safety assessment. Using a pre-trained safety scoring model, the safety score of each physical region is determined based on the temporal features, enabling a quantitative assessment of the safety status of physical regions. A target score for each physical region is determined based on the semantic features, temporal features, and safety score. The parking area is then determined from multiple physical regions based on the target score. This multi-dimensional feature-based target score determination makes the determined target score more accurate, significantly improving the accuracy and rationality of parking area selection, effectively reducing safety risks during parking, and improving the efficiency and safety of vehicle parking.
Smart Images

Figure CN121536283B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of vehicle control technology, and in particular to a parking control method, a storage medium, and a vehicle. Background Technology
[0002] With the development of the automotive industry, finding suitable parking areas has become a major challenge for drivers. Currently, parking area identification is often based on single sensory data, which can lead to inaccurate identification due to a lack of comprehensive understanding of the vehicle's surrounding environment.
[0003] In view of this, how to fully understand the environment around the vehicle to accurately identify the parking area has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of this disclosure is to provide a parking control method, storage medium and vehicle to solve the problem of inaccurate parking area identification caused by the inability to fully understand the vehicle's surrounding environment in the prior art.
[0005] To achieve the above objectives, the first aspect of this disclosure provides a parking control method, the method comprising: The system acquires environmental data surrounding the vehicle and the vehicle's own state data, and constructs a semantic graph based on the environmental data and the state data; wherein the semantic graph includes multiple physical regions; Using a pre-trained cross-modal attention network, the node features corresponding to multiple nodes in the semantic graph are fused to obtain the semantic features of each physical region; Using a pre-trained spatiotemporal graph generation model, the temporal characteristics of each physical region are determined based on the semantic features; Using a pre-trained security scoring model, the security score for each physical region is determined based on the temporal characteristics. The target score for each physical region is determined based on the semantic features, the temporal features, and the security score, and the parking area is determined from multiple physical regions based on the target score.
[0006] Based on the same inventive concept, a second aspect of this disclosure provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the method described above.
[0007] Based on the same inventive concept, a third aspect of this disclosure proposes a vehicle that includes the storage medium described in the second aspect.
[0008] As described above, the parking control method, storage medium, and vehicle disclosed herein acquire environmental data surrounding the vehicle and the vehicle's own state data, and construct a semantic graph based on the environmental and state data. Utilizing a pre-trained cross-modal attention network, the semantic features of each physical region are obtained by fusing the node features corresponding to multiple nodes in the semantic graph, effectively integrating multi-source environmental and state data and enhancing the understanding of complex environments. Using a pre-trained spatiotemporal graph generation model, the temporal features of each physical region are determined based on the semantic features, capturing the dynamic changes of physical regions at different times and providing a temporal dimension basis for safety assessment. Using a pre-trained safety scoring model, the safety score of each physical region is determined based on the temporal features, enabling a quantitative assessment of the safety status of physical regions. A target score for each physical region is determined based on the semantic features, temporal features, and safety score. The parking area is then determined from multiple physical regions based on the target score. This multi-dimensional feature-based target score determination makes the determined target score more accurate, significantly improving the accuracy and rationality of parking area selection, effectively reducing safety risks during parking, and improving the efficiency and safety of vehicle parking. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart of a parking control method according to an embodiment of the present disclosure; Figure 2 This is a flowchart of a parking area identification method based on multimodal data according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the parking control device according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0012] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0013] Based on the background description, current methods for identifying safe parking areas primarily rely on single-modal perception data, such as images or LiDAR, and employ rule-based templates or traditional detection models for area selection. This makes them ill-suited to complex environments, dynamically changing targets, and diverse user preferences in real-world road scenarios. On one hand, images are easily affected by occlusion and lighting conditions, leading to inaccurate parking area identification. On the other hand, the lack of ability to model the correlations between multimodal information limits the system's understanding of potential parking areas to a static perspective, making it unable to predict future changes in the parking area's state. Furthermore, current methods generally lack user intent understanding mechanisms, resulting in a lack of personalized parking area recommendations that fail to meet the dual requirements of advanced driver assistance systems (ADAS) for both safety and user experience.
[0014] As mentioned above, how to comprehensively understand the vehicle's surrounding environment to accurately identify parking areas has become an important research question.
[0015] Based on the above description, such as Figure 1 As shown, the parking control method proposed in this embodiment includes: Step 101: Obtain environmental data around the vehicle and the vehicle's own state data, and construct a semantic graph based on the environmental data and the state data; wherein, the semantic graph includes multiple physical regions.
[0016] In practice, during vehicle operation, onboard sensors acquire multimodal environmental perception data in real time to support subsequent intelligent parking area recognition tasks. This multimodal environmental perception data includes both the surrounding environment and the vehicle's own state data.
[0017] Multiple nodes and edges between every two nodes are determined based on environmental and state data. A semantic graph is then constructed based on these nodes and edges. The semantic graph is a graphical structure based on nodes and edges, used to represent entities, concepts, and the relationships between them. Nodes represent specific entities (e.g., image nodes, text nodes, physical nodes, map nodes), and edges represent the associations between entities (e.g., spatial similarity, semantic relevance, temporal similarity).
[0018] Step 102: Using a pre-trained cross-modal attention network, the node features corresponding to multiple nodes in the semantic graph are fused to obtain the semantic features of each physical region.
[0019] In practice, cross-modal attention networks are used to perform efficient information fusion and node feature enhancement on the constructed semantic graph.
[0020] A pre-trained cross-modal attention network is used to map image node features, text node features, physical node features, and map node features to a predefined semantic space to obtain standard node features. Neighboring node features are determined for each standard node feature, and the standard node features and neighboring node features are fused to obtain updated node features. These updated node features are then used as the semantic features of the physical region corresponding to the standard node features.
[0021] Step 103: Using a pre-trained spatiotemporal graph generation model, determine the temporal characteristics of each physical region based on the semantic features.
[0022] In practice, a spatiotemporal graph generation model is used to construct high-dimensional dependencies between cross-frame semantic graphs, forming a spatiotemporal semantic evolution modeling module to achieve dynamic tracking and evolutionary identification of potential parking areas.
[0023] Using a pre-trained spatiotemporal graph generation model, nodes of the same type are identified from the semantic features of multiple consecutive frames, and temporal node features of each physical region are determined based on these nodes. Using the same-type nodes, spatial node features of each physical region are determined from the semantic features of multiple consecutive frames. Temporal features of each physical region are determined based on both temporal and spatial node features.
[0024] Step 104: Using a pre-trained security scoring model, determine the security score for each physical region based on the temporal characteristics.
[0025] In practice, in order to achieve accurate and secure assessment of candidate physical regions, a highly reliable scoring mechanism needs to be established to quantify the nodes processed by the cross-modal attention network and the spatiotemporal graph generation model.
[0026] The pre-trained security scoring model can be a sparse Gaussian process regression model. The sparse Gaussian process regression model is used to determine the security score for each physical region based on temporal characteristics.
[0027] Step 105: Determine the target score for each physical region based on the semantic features, the temporal features, and the security score; determine the parking area from multiple physical regions based on the target score.
[0028] In practice, semantic features, temporal features, and safety scores are fused to obtain the target score for each physical region. The parking region with the highest target score is then determined from multiple physical regions. Thus, the parking region is determined based on three dimensions: semantic features, temporal features, and safety scores, ensuring that the parking region takes into account both visual semantic performance and temporal stability.
[0029] This disclosure achieves a personalized intelligent identification and rating recommendation mechanism for safe areas in dynamic environments by integrating a cross-modal attention network, a spatiotemporal graph generation model, and a sparse Gaussian process regression model to construct a multimodal semantic graph and incorporating user preference embedding. This disclosure addresses the problems of reliance on single-modal perception, lack of dynamic state modeling, and user preference guidance in related technologies for area identification, thus improving the accuracy, stability, and personalized recommendation capabilities of parking area identification.
[0030] Through the above embodiments, environmental data surrounding the vehicle and the vehicle's own state data are acquired, and a semantic graph is constructed based on the environmental and state data. Using a pre-trained cross-modal attention network, the node features corresponding to multiple nodes in the semantic graph are fused to obtain the semantic features of each physical region. This effectively integrates multi-source environmental and state data, enhancing the understanding of complex environments. Using a pre-trained spatiotemporal graph generation model, the temporal features of each physical region are determined based on the semantic features, capturing the dynamic changes of physical regions at different times and providing a temporal dimension for safety assessment. Using a pre-trained safety scoring model, the safety score of each physical region is determined based on the temporal features, enabling a quantitative assessment of the safety status of physical regions. A target score for each physical region is determined based on semantic features, temporal features, and safety scores. Parking areas are then determined from multiple physical regions based on the target scores. This multi-dimensional feature-based approach makes the target scores more accurate, significantly improving the accuracy and rationality of parking area selection, effectively reducing safety risks during parking, and enhancing the efficiency and safety of vehicle parking.
[0031] In some embodiments, step 101 includes: Step 1011: Obtain environmental data around the vehicle and the vehicle's own status data; wherein, the environmental data includes image data and point cloud data, and the status data includes map data.
[0032] In practice, during vehicle operation, onboard sensors acquire multimodal environmental perception data in real time to support subsequent intelligent parking area recognition tasks. This multimodal environmental perception data includes both the surrounding environment and the vehicle's own state data.
[0033] The vehicle-mounted sensors include: a forward-facing camera, a LiDAR point cloud sensor, and a positioning module sensor. The positioning module sensor includes a Global Positioning System (GPS) and an Inertial Measurement Unit (IMU).
[0034] Multimodal environmental perception data includes: image data, point cloud data, map data, pose data, and user preference data. Specifically, image data is acquired using a forward-looking camera, point cloud data is acquired using LiDAR, map data is acquired using GPS in the positioning module sensor, and pose data is acquired using IMU in the positioning module sensor.
[0035] Step 1012: Perform time synchronization processing on the image data, the point cloud data, and the map data to obtain time synchronization data, and perform standardization processing on the time synchronization data to obtain standard image data, standard point cloud data, and standard map data.
[0036] In practice, the multimodal environment perception data all originate from the vehicle's original perception system and human-machine interaction system, possessing advantages in stability and structure. To ensure the uniformity of subsequent multimodal graph structure construction, the multimodal environment perception data needs to undergo time synchronization processing, standardization processing, and semantic vector encoding processing to form unified standard data.
[0037] The image data consists of RGB images (red, green, blue) captured by an onboard camera. The image resolution is 1920x1080, and the frame rate is 30 frames per second. These images are primarily used to extract visual semantic features and assist in spatial structure recognition. Before entering the subsequent network, the image data undergoes preprocessing steps such as distortion correction, size normalization, and color space conversion to obtain standard image data. A pre-trained image encoder using a multimodal alignment model extracts semantic features from the standard image data, outputting a one-dimensional fixed-length image embedding vector. This image embedding vector serves as the initial semantic representation for the image nodes.
[0038] Point cloud data was acquired using a forward-mounted 64-line LiDAR with a sampling frequency of 10Hz. This point cloud data was used to obtain 3D environmental spatial information. After acquisition, the point cloud data underwent point cloud filtering, followed by voxel grid downsampling to compress the data volume, ultimately forming a sparse but representative spatial structure representation. The point cloud filtering process included ground separation, noise removal, and cluster pre-segmentation. After location normalization and spatial coordinate transformation, each cluster of point cloud data was treated as a potential spatial region or obstacle unit, used as a physical node in the subsequent semantic graph.
[0039] The map data is acquired using an in-vehicle high-precision map system and includes at least one of the following layers: road structure, lane topology, static obstacles, no-stopping zones, service areas, and signs and markings. During map data acquisition, high-precision map tiles of the matching area are retrieved based on vehicle positioning information, and structured labels are extracted from these tiles as the initial semantic descriptions of map nodes. The structured labels include at least one of the following types: area usage, legal constraints, and road boundaries. The structured labels are then converted into vectorized embeddings and added to the semantic map.
[0040] The vehicle's pose data is acquired using a fusion positioning module combining onboard GPS and IMU. The pose data includes at least one of the following: latitude and longitude, heading angle, and velocity. This pose data provides a unified spatial reference frame for each node in the semantic graph. Within the semantic graph, each frame of multimodal environmental perception data is bound to the vehicle's current position and pose data, ensuring the spatiotemporal alignment between multiple frames of multimodal environmental perception data.
[0041] User preference data is provided by the navigation system, voice interaction module, or mobile terminal input. User preference data includes at least one of the following: the user's natural language needs, historical selection preferences, and specific scenario instructions. The user preference data is input in natural language form and processed by the language encoding module to generate a fixed-dimensional semantic embedding vector. This vector serves as the user preference text node or edge attribute, influencing the attention weights in the semantic graph and subsequent scoring mechanisms.
[0042] Step 1013: Construct a semantic graph based on the standard image data, the standard point cloud data, and the standard map data.
[0043] In practice, to achieve a unified input format for multimodal environment perception data, all processed standard perception data is encapsulated into a unified data packet. This packet includes standard image data, standard point cloud data, standard map data, standard pose data, and standard user preference data, all bound with a unified timestamp. This data packet serves as the input foundation for semantic graph construction, ensuring accurate spatial and semantic alignment between different modalities. After the multimodal environment perception data acquisition and preprocessing process is completed, the standard data required for semantic graph construction is formed.
[0044] The above scheme acquires environmental data surrounding the vehicle and the vehicle's own status data. Environmental data includes image data and point cloud data, while status data includes map data. Time synchronization processing is performed on the image data, point cloud data, and map data to obtain time-synchronized data. Standardization processing is then applied to this time-synchronized data to obtain standard image data, standard point cloud data, and standard map data. This achieves time synchronization and standardization of the image data, point cloud data, and map data, ensuring that the resulting standard image data, standard point cloud data, and standard map data are standardized data with unified time and format. A semantic graph is constructed based on the standard image data, standard point cloud data, and standard map data, making the resulting semantic graph more accurate and enabling the fusion of multiple standard data sets.
[0045] In some embodiments, step 1014 includes: Step 1014A: Encode the standard image data to obtain an image embedding vector, and use the image embedding vector as an image node.
[0046] In practice, after completing the collection and standardization of multimodal environmental perception data, the semantic graph construction phase begins. The semantic graph, as the core intermediate representation, carries multimodal environmental perception data from images, point clouds, high-precision maps, and user input, aiming to model the regional features and their interrelationships in the vehicle's surrounding environment in a unified manner. The construction of the semantic graph not only requires semantic alignment between modalities but also needs to consider multi-dimensional relationships such as spatial location, physical topology, and user intent, possessing strong expressiveness and scalability to provide complete input for subsequent graph neural networks and temporal modeling. Each node in the semantic graph represents a scene element, and nodes are mainly divided into four categories: image nodes, text nodes, physical nodes, and map nodes.
[0047] Image nodes are determined based on standard image data, using the image embedding vector generated by the image encoder as the initial semantic representation of the image node. Image nodes are used to identify the visual environment observed from the vehicle's current perspective. The visual environment includes static areas such as buildings, open spaces, and green spaces, as well as dynamic elements such as pedestrians and vehicles.
[0048] Step 1014B: Extract text embedding vectors from the standard map data and use the text embedding vectors as text nodes.
[0049] In practice, text nodes are determined based on linguistic information from standard user preference data and standard map data. A natural language processing module extracts text embedding vectors from these data and uses them as text nodes. These text nodes provide the semantic graph with the ability to actively express intent, enabling the model to offer personalized recommendations.
[0050] Step 1014C: Determine the point cloud cluster center and cluster feature vector based on the standard point cloud data, determine the physical region according to the point cloud cluster center and the cluster feature vector, and use the physical region as a physical node.
[0051] In practice, physical nodes are primarily determined based on standard point cloud data, representing physical regions of 3D entities in the environment. Each point cloud data point represents a structurally stable spatial region or obstacle unit, and the geometric location and type attributes of the physical region are defined by the coordinates of its center point and clustering feature vectors. Physical nodes enhance the modeling capabilities for spatial accessibility, occlusion relationships, and region boundaries.
[0052] Step 1014D: Extract structured labels from the standard map data and map the structured labels to map nodes.
[0053] In practice, map nodes are determined based on structured labels extracted from standard map data. These structured labels include service areas, no-parking zones, and emergency avoidance zones. Each type of structured label is mapped to a corresponding map node, and a semantic vector is attached for use in fusion calculations.
[0054] Step 1014E: Determine the spatial similarity, semantic relevance, and temporal similarity between every two nodes, and determine the edge between every two nodes based on the spatial similarity, semantic relevance, and temporal similarity.
[0055] In practice, the nodes are connected by edges to form a complete semantic graph. The principles for setting edges include three aspects: spatial similarity, semantic relevance, and temporal similarity.
[0056] In this system, image nodes and physical nodes are connected by edges based on spatial similarity (e.g., spatial location projection relationships). For example, an edge is generated between an empty area detected in an image and its corresponding planar clustering unit in the point cloud. Text nodes are connected to image nodes and map nodes based on semantic relevance. Semantic relevance is calculated using a pre-trained model of language and images. When the semantic relevance is greater than a preset relevance threshold, a semantic edge is established to convey user preference data. Physical nodes are connected by edges based on spatial similarity (e.g., geometric distance and co-occurrence frequency of historical trajectories) to enhance the modeling of traversability and historical security relevance between spatial regions in the scene. Map nodes are connected by edges with other nodes based on spatial similarity (e.g., layer overlap). For example, if an empty area in an image is located within a service area marked on a map, an edge is established between the corresponding image node and the service area map node.
[0057] Each edge in the semantic graph is assigned a set of weight attributes for subsequent attention calculation. The weight attributes include at least one of the following: spatial distance, semantic similarity, historical risk level, and map priority.
[0058] Step 1014F: Construct a semantic graph based on the image nodes, text nodes, physical nodes, map nodes, and the edges between every two nodes.
[0059] In practice, all nodes and edges form a directed weighted semantic graph, which serves as the input to the subsequent graph neural network. To ensure the stability and real-time performance of the system under high-frequency input, the semantic graph construction module adopts a sliding window mechanism, constructing only a data subgraph covering the current frame and several frames before and after at each time step, and dynamically updating the node and edge set as the vehicle moves, thus maintaining the sparsity and information density of the semantic graph.
[0060] The semantic graph includes a set of nodes, a set of edges, and a node feature matrix, which are used to input the semantic graph into a cross-modal attention network, supporting interactive fusion and semantic aggregation processing between different modalities. By constructing the semantic graph, the system ensures that vehicles have structured semantic representation capabilities in complex perception environments, which is the foundation of the entire parking area recognition process.
[0061] The above scheme encodes standard image data to obtain image embedding vectors, which are then used as image nodes. Text embedding vectors are extracted from standard map data, and these are used as text nodes. Point cloud cluster centers and cluster feature vectors are determined based on standard point cloud data, and physical regions are identified based on these cluster centers and feature vectors, with these physical regions used as physical nodes. Structured labels are extracted from standard map data and mapped to map nodes. Spatial similarity, semantic relevance, and temporal similarity between any two nodes are determined, and edges between these pairs are established. A semantic graph is constructed based on image nodes, text nodes, physical nodes, map nodes, and the edges between each pair of nodes. This constructed semantic graph integrates standard image data, standard point cloud data, and standard map data, effectively consolidating multi-source environmental and state data and enhancing the understanding of complex environments.
[0062] In some embodiments, step 102 includes: Step 1021: Using a pre-trained cross-modal attention network, image node features, text node features, physical node features, and map node features are mapped to a preset semantic space to obtain standard node features.
[0063] In specific implementation, the cross-modal attention network is one of the core modules of this disclosure embodiment. The cross-modal attention network is used for efficient information fusion and node feature enhancement on the constructed semantic graph. The input to the cross-modal attention network is a pre-generated semantic graph. Combining the semantic and structural relationships between multimodal features, it utilizes graph neural network mechanisms to perform context enhancement and global perception processing on nodes in the semantic graph. Unlike traditional graph convolutional structures, this graph neural network introduces a cross-modal attention mechanism, which can dynamically adjust the information transmission intensity between different modalities, especially establishing an explicit coupling relationship between user semantic preferences and environmental perception data, thereby improving the semantic recognition capability of parking areas.
[0064] The input to the cross-modal attention network is a semantic graph output by the graph construction module. This semantic graph includes a set of nodes, a set of edges, and an initial feature vector for each node. Image node features are image embedding vectors extracted by the image encoder, text node features are text embedding vectors extracted by the language model, physical node features contain spatial structure encoding of point cloud clusters, and map node features consist of structured labels from a high-precision map. Each edge in the semantic graph, in addition to its connectivity, contains multi-dimensional edge attributes, including spatial distance, similarity, and map constraint weights, used to adjust the intensity of information propagation between different nodes.
[0065] In the structural design of the cross-modal attention network, a multi-head graph neural network based on an attention mechanism is adopted to weighted aggregate information between nodes of different modalities. During processing, each node not only considers the features of its corresponding first-order neighbor nodes, but also introduces edge attributes as auxiliary adjustment factors to dynamically adjust the direction and intensity of information propagation. When image nodes aggregate text nodes, they automatically adjust the attention distribution according to the semantic alignment between text and image data to achieve explicit injection of user intent. When physical nodes interact with map nodes, they suppress or amplify the feature enhancement process according to regional regulations to highlight the feature expression of legal and safe areas. Text nodes play a guiding role in the entire graph. The features of text nodes not only propagate to surrounding related nodes, but also affect the overall feature shift direction of the network, realizing semantically driven semantic graph reconstruction.
[0066] To enhance the adaptability to multimodal environment perception data in complex scenarios, the cross-modal attention network adopts a unified cross-modal embedding strategy, mapping all modal features to the same semantic space and processing them uniformly through a shared attention mechanism. This strategy eliminates the semantic gap between multimodal features, enabling direct comparison and fusion of image node features, point cloud node features, and text node features within the same embedding space.
[0067] The pre-training process of the cross-modal attention network includes: supervised learning based on historical labeled data from automotive manufacturers, with the goal of maximizing the feature response values of nodes in parking areas. A multi-task loss function is introduced during training to consider both the prediction accuracy of safety scores and the optimization of attention matching between different modalities in the semantic graph, thereby improving the overall fusion effect. To adapt to in-vehicle computing resources, the parameters of the cross-modal attention network are pruned and quantized for compression, ensuring its real-time inference capabilities on embedded platforms.
[0068] Step 1022: Determine the neighbor node features of each standard node feature, fuse the standard node features and the neighbor node features to obtain updated node features, and use the updated node features as the semantic features of the physical region corresponding to the standard node features.
[0069] In practice, after each round of propagation, the semantic features of each node will be updated to the weighted fusion result of the aggregated features of the corresponding node and its neighboring nodes. The vector of updated node features will be used for graph time-series modeling in subsequent steps.
[0070] By extracting semantic features from each physical region, the semantic representation of each node in the semantic graph is enhanced. Particularly in semantically ambiguous regions or areas with severe structural occlusion, the cross-modal attention network can significantly improve the robustness and accuracy of recognition through cross-completion of multimodal data. The final output is a set of structured graph vectors, where the updated node features of each node are an updated multimodal fusion feature vector, serving as the foundational input for subsequent temporal modeling and region scoring prediction. The cross-modal attention network achieves the crucial transformation from raw perceptual data to semantically enhanced representation, constructing the core cognitive capability of the safe zone recognition system.
[0071] The above scheme utilizes a pre-trained cross-modal attention network to map image node features, text node features, physical node features, and map node features to a preset semantic space to obtain standard node features. This eliminates the semantic gap between multimodal data, enabling direct comparison and fusion of image node features, point cloud node features, and text node features within the same embedding space. The neighboring node features of each standard node feature are determined, and the standard node features and neighboring node features are fused to obtain updated node features. These updated node features are then used as the semantic features of the physical region corresponding to the standard node features. This results in the updated node features for each node being a weighted fused node feature, enabling cross-completion of multimodal data in semantically ambiguous regions or regions with severe structural occlusion, significantly improving the robustness and accuracy of semantic feature recognition in physical regions.
[0072] In some embodiments, step 103 includes: Step 1031: Using a pre-trained spatiotemporal graph generation model, determine nodes of the same type from the semantic features of multiple consecutive frames, and determine the temporal node features of each physical region based on the nodes of the same type.
[0073] In practical implementation, after the cross-modal attention network completes the feature fusion of each node in the single-frame semantic map, it is necessary to further model the multi-frame semantic map in continuous temporal sequence to capture the evolution of physical regions over time. During dynamic vehicle operation, environmental and state data change continuously. A safe parking area in one frame may become unavailable in the next frame due to obstruction, approaching targets, or changes in right-of-way. Therefore, relying solely on static semantic maps is insufficient to support the judgment of the stability and continued accessibility of parking areas. To address this, a spatiotemporal graph generation model (based on a graph-structured spatiotemporal Transformer model) is introduced to construct high-dimensional dependencies between cross-frame semantic maps, forming a spatiotemporal semantic evolution modeling module to achieve dynamic tracking and evolutionary identification of potential parking areas.
[0074] The spatiotemporal graph generation model takes as input the semantic features of multiple consecutive frames output by a cross-modal attention network. The length of the time window can be customized; for example, setting the time window to six forward frames with a 100-millisecond interval between each frame covers a total time span of 600 milliseconds. Each frame of the semantic graph contains a set of nodes and corresponding updated node features. The system first pairs nodes of the same type in the semantic features. For example, empty areas that are close in location in multiple consecutive frames are considered as the evolutionary state of the same physical entity at different times, thus forming temporal node trajectory lines. These temporal node trajectory lines serve as input units, constituting the basic elements of the temporal features (spatiotemporal graph structure).
[0075] Step 1032: Using a pre-trained spatiotemporal graph generation model, determine the spatial node features of each physical region from the semantic features of multiple consecutive frames.
[0076] In practical implementation, the spatiotemporal graph generation model internally employs a graph-enhanced temporal Transformer model. The basic modules of this model are stacked self-attention encoders. Unlike the standard Transformer model, which only processes sequence vectors, each time node in the graph-enhanced temporal Transformer model is itself a graph structure. The attention mechanism aggregates the evolution states of similar nodes from different frames in the temporal dimension, while maintaining the topological invariance of the graph structure within the spatial dimension. Temporal attention is used to discover trends in node state changes, while spatial attention strengthens the stable expression of semantic dependencies between nodes. When processing the trajectory of each node, the spatiotemporal graph generation model not only considers the node's own state changes but also integrates historical change information from neighboring nodes, thereby obtaining a safe trend prediction capability for future moments.
[0077] Maintaining the connections between nodes is accomplished through a trajectory consistency constraint mechanism. In real-world scenarios, due to occlusion, radar noise, and other factors, the same physical region may be lost in detection across some frames. To compensate for this gap, the system employs a joint matching approach combining feature similarity and pose prediction. This method completes the trajectories of nodes with similar structures and reasonable positions in consecutive frames, ensuring the continuity of the time series. Simultaneously, a time decay factor is introduced, causing the weights of features from nodes far from the current frame to gradually decrease during aggregation, ensuring the model's response to the current scene is more sensitive.
[0078] Step 1033: Determine the temporal characteristics of each physical region based on the temporal node characteristics and the spatial node characteristics.
[0079] In practice, the temporal characteristics of each physical region are determined based on the temporal and spatial characteristics of the nodes. Temporal attention is used to discover trends in node state changes, while spatial attention reinforces the stable expression of semantic dependencies between nodes.
[0080] The pre-training process of the spatiotemporal map generation model involves classifying changes in safety status for each node trajectory using regional state changes as labels. Samples are derived from historical driving data, including records of successful stops, risk warnings, and interference events in various regions. These records are then clustered and labeled using actual map tags to create a positive and negative sample library. During training, the spatiotemporal map generation model learns which spatial regions possess continuous safety and which exhibit sudden instability, thereby improving its ability to perceive the continuous availability of parking areas.
[0081] After the spatiotemporal map generation model is trained, it can output a stability score for each spatial region in the next few hundred milliseconds, which serves as an important input indicator for parking area evaluation.
[0082] The spatiotemporal graph generation model effectively bridges the gap between static perception and dynamic semantic reasoning, enabling the system to understand whether a currently accessible parking area maintains continuous safety. It also provides dynamic prior conditions for the final score regression and safety recommendation. The inference output of the spatiotemporal graph generation model is a state prediction vector for each candidate region node at the current and several future time steps. This state prediction vector includes the safety trend distribution and change confidence. This vector, along with semantic graph features and user preferences, participates in the subsequent region scoring calculation process, completing the temporal modeling portion of the entire safe zone identification chain.
[0083] The above scheme utilizes a pre-trained spatiotemporal graph generation model to identify nodes of the same type from semantic features across multiple consecutive frames, and then determines the temporal node features of each physical region based on these nodes. The pre-trained spatiotemporal graph generation model also determines the spatial node features of each physical region from semantic features across multiple consecutive frames. Finally, the temporal features of each physical region are determined based on both temporal and spatial node features. In this way, the temporal features of each physical region can comprehensively reveal the state change trends and stability expressions of the physical region, thereby determining whether the physical region possesses sustainable security.
[0084] In some embodiments, step 104 includes: Step 1041: Using a pre-trained security scoring model, determine the security score and confidence level of each physical region based on the time-series characteristics.
[0085] In practical implementation, to achieve accurate and safe assessment of candidate physical regions, a highly reliable scoring mechanism needs to be established to quantify the nodes processed by the cross-modal attention network and spatiotemporal graph generation model. The key to this scoring mechanism is assigning a safety score to each physical node, characterizing the feasibility of the physical region corresponding to that node as a parking area during the current and predicted time periods. The score should be sensitive to dynamic environmental factors, adaptable to user preferences, and have a reasonable expression of uncertainty to support downstream decision-making modules in making interpretable recommendations. Therefore, this step introduces a sparse Gaussian process regression model for predicting regional safety in the multimodal embedding space.
[0086] The pre-training process of the sparse Gaussian process regression model is as follows: Historical operational data accumulated by automakers is acquired and used as sample data. This historical operational data includes a large amount of manually labeled and automatically recorded safety event data. The historical operational data originates from observations and behavioral feedback of the environment during vehicle operation. For example, it includes at least one of the following: historical stop success rate, frequency of regional accidents, target approach frequency, navigation deviation records, and passenger-triggered behaviors (e.g., early stop requests). After spatial alignment, the historical operational data is mapped to node positions in the semantic graph to generate labeled samples. Each labeled sample is bound to a spatial region embedding vector and a corresponding safety score. The score values are typically normalized to between 0 and 1, with higher scores indicating greater safety. The training set includes a wide distribution of positive and negative samples to ensure that the sparse Gaussian process regression model can accurately fit the distribution characteristics of high-risk, boundary, and ideal regions.
[0087] The sparse Gaussian process regression model is trained using a hybrid labeled dataset independently collected by the automaker. During the input feature preprocessing stage, a feature selection mechanism is introduced to remove redundant information, retaining only modality fusion dimensions strongly correlated with the scoring, such as visual texture sparsity, region occlusion probability, historical navigation traversal rate, and user target overlap. A kernel function adjustment strategy is incorporated into the structural design of the sparse Gaussian process regression model, enabling it to adaptively select a more suitable fitting surface shape based on the data distribution in different physical regions during training, thereby more accurately reflecting the risk distribution of the actual physical region.
[0088] The sparse Gaussian process regression model takes semantic and temporal features from a cross-modal attention network and a spatiotemporal graph generation model as input, and outputs a safety score in continuous value form. The model uses a sparse Gaussian process as the regressor, enabling it to reasonably avoid uncertain physical regions. Compared to conventional neural network regression methods, the sparse Gaussian process regression model maintains controllable output distribution even when faced with scarce boundary samples or extreme variations in regional distribution, avoiding distortion from extreme scores. The sparsity mechanism allows the sparse Gaussian process regression model to make effective predictions using a limited set of representative samples in a high-dimensional space, reducing computational complexity and adapting to resource constraints in vehicle environments.
[0089] The output of the sparse Gaussian process regression model includes not only the safety score but also the prediction variance, which is an estimate of the uncertainty of the safety score for a physical region. In practical applications, this uncertainty estimate can be used to exclude high-risk predictions and improve the stability of the recommendation system. For physical regions with safety scores close to the boundary but high uncertainty, the system will automatically reduce the priority of the corresponding physical regions in subsequent fusion stages, thereby achieving risk control for sparse data regions.
[0090] Step 1042: Sort the multiple physical regions according to the security score and the confidence level to obtain the priority of the multiple physical regions.
[0091] In practice, after training, the sparse Gaussian process regression model is deployed as a service component, running in parallel with the graph inference module and the user preference guidance module. During each inference process, the system calls the regressor on all candidate physical region nodes in the current semantic graph to generate safety scores and confidence levels, which are used to construct a priority list of multiple physical regions. The priorities of these multiple physical regions directly participate in the subsequent multi-model fusion and ranking decision-making process, becoming a crucial foundation for the final parking area recommendation result. The scoring mechanism realizes the transformation from semantic understanding to numerical evaluation, providing decision-level quantitative support capabilities for the entire intelligent safe area identification system.
[0092] The above scheme utilizes a pre-trained safety scoring model to determine the safety score and confidence level of each physical area based on time-series characteristics. This allows for a rapid and accurate assessment of the safety status of each physical area, enabling a quantitative evaluation of its safety condition. Multiple physical areas are then prioritized based on their safety scores and confidence levels, which is used to subsequently determine parking areas.
[0093] In some embodiments, step 105 includes: Step 1051: Map the semantic features and the temporal features to a preset space to obtain standard semantic features and standard temporal features, and perform weighted fusion processing on the standard semantic features and the standard temporal features to obtain the region representation of each physical region.
[0094] In practice, the distribution of safety scores for physical areas enables probabilistic prediction and uncertainty estimation of future area states. After completing the regression prediction of safety scores for physical areas, the outputs of multiple models need to be fused to generate a more robust final safety evaluation that better meets actual driving needs. Since cross-modal attention networks are used for semantic information fusion between modalities, spatiotemporal graph generation models are used to extract the evolutionary trends of physical areas, and sparse Gaussian process regression models provide safety scores for risk quantification prediction, the three models focus on different dimensions. During fusion, not only must their respective weights be considered, but also potential semantic conflicts and spatial distribution biases between the outputs of different models must be addressed. Therefore, this step designs a multi-model fusion mechanism, combined with a reverse consistency optimization strategy, to effectively integrate different information sources and improve the stability and consistency of the recommended parking areas.
[0095] The fusion strategy uses physical nodes of a physical region as the basic unit. For each candidate physical region's regional node, the outputs of the three models are collected separately. The cross-modal attention network outputs the globally fused semantic features of the physical node, reflecting the semantic features of the physical region in the current environment and its contextual association; the spatiotemporal graph generation model outputs the temporal features of the physical node, reflecting the state stability evolution characteristics of the physical region over future frames; and the sparse Gaussian process regression model provides the safety score and prediction variance of the physical node. The outputs of the three models form a fused information set, which is used to calculate the target score of the corresponding physical node.
[0096] During the fusion process, modal features are first standardized to ensure consistency in dimensionality, scale, and confidence level among the outputs of the three models. Semantic and temporal features are both mapped to a unified space and combined using similarity weighting to form a regional representation of the physical region.
[0097] Step 1052: Determine the semantic score of the standard semantic feature and the temporal score of the standard temporal feature in the region representation, determine the first score deviation between the semantic score and the security score, and determine the second score deviation between the temporal score and the security score.
[0098] In practical implementation, a reverse consistency optimization mechanism is introduced. Specifically, after the initial fusion results are generated, the system uses a lightweight reverse verification network to perform semantic consistency verification on high-scoring physical regions. Based on a simplified dual-path graph attention structure, the predicted results of physical regions are compared in reverse, using semantic features and temporal features as inputs respectively. If the predicted physical region has a significant scoring deviation in one modality, the original fusion weights are adjusted through a feedback mechanism to update the target score of the physical region.
[0099] Step 1053: In response to determining that both the first scoring deviation and the second scoring deviation are less than a preset scoring threshold, a target score for each physical region is determined based on a preset semantic weight, a preset temporal weight, and a preset security weight.
[0100] In practice, when both the first scoring deviation and the second scoring deviation are less than the preset scoring threshold, it means that there is no significant scoring deviation in semantic scoring and temporal scoring. There is no need to adjust the weights corresponding to each score. The target score for each physical region is determined directly based on the preset semantic weight, preset temporal weight, and preset security weight.
[0101] Step 1054: In response to determining that the first scoring deviation is greater than or equal to a preset deviation threshold and the second scoring deviation is less than a preset scoring threshold, a target score for each physical region is determined based on a first updated semantic weight, a first updated temporal weight, and a preset security weight; wherein the first updated semantic weight is less than the preset semantic weight, and the first updated temporal weight is greater than the preset temporal weight.
[0102] In practice, when the first scoring deviation is greater than or equal to the preset deviation threshold and the second scoring deviation is less than the preset scoring threshold, it indicates that there is a significant scoring deviation in the semantic score. It is necessary to reduce the semantic weight corresponding to the semantic score and increase the temporal weight corresponding to the temporal score. The target score for each physical region is determined based on the first updated semantic weight, the first updated temporal weight and the preset security weight.
[0103] For example, the preset semantic weight is 0.3, the preset temporal weight is 0.3, and the preset security weight is 0.4. When the first scoring deviation is greater than or equal to the preset deviation threshold and the second scoring deviation is less than the preset scoring threshold, the first updated semantic weight is 0.2 and the first updated temporal weight is 0.4.
[0104] Step 1055: In response to determining that the second scoring deviation is greater than or equal to a preset deviation threshold and the first scoring deviation is less than a preset scoring threshold, a target score for each physical region is determined based on a second updated semantic weight, a second updated temporal weight, and a preset security weight; wherein the second updated semantic weight is greater than the preset semantic weight, and the second updated temporal weight is less than the preset temporal weight.
[0105] In practice, when the second score deviation is greater than or equal to the preset deviation threshold and the first score deviation is less than the preset score threshold, it indicates that there is a significant score deviation in the time-series score. It is necessary to reduce the semantic weight corresponding to the time-series score and increase the time-series weight corresponding to the semantic score. The target score for each physical region is determined based on the second updated semantic weight, the second updated time-series weight and the preset security weight.
[0106] For example, the preset semantic weight is 0.3, the preset temporal weight is 0.3, and the preset security weight is 0.4. When the second scoring deviation is greater than or equal to the preset deviation threshold and the first scoring deviation is less than the preset scoring threshold, the second updated semantic weight is 0.4, and the second updated temporal weight is 0.2.
[0107] The aforementioned optimization strategy effectively suppresses high-confidence misclassifications caused by local model overfitting. For example, if a physical region performs well visually and semantically, but its future state changes rapidly and its risk is high, the system can correct the overestimation of the physical region's score through temporal features, ensuring temporal consistency of the output results. Similarly, if a physical region performs stably over time, but there is occlusion or misidentification in the visual modality, the high score of the physical region can also be reasonably suppressed. Based on the final fused target score, a list of ranked scores with confidence is generated for all candidate physical regions for use by downstream modules.
[0108] Furthermore, to ensure the robustness and adaptability of the fusion mechanism, the fusion module supports a configurable weight allocation strategy. Weights can be set at fixed ratios, such as a preset semantic weight of 40%, a preset temporal weight of 30%, and a preset safety weight of 30%, or they can be dynamically adjusted at runtime based on data quality. For example, in cases of severe visual occlusion or drastic changes in lighting, the weight of semantic features in the fusion process will be reduced, while the reliance on temporal features will be increased; conversely, in cases of strong dynamic interference or unstable trajectory prediction, semantic features and safety scores will be prioritized. This adaptive mechanism enhances responsiveness and stability in different scenarios.
[0109] The fused target score will serve as the final recommendation basis for selecting a set of optimal parking areas. It can also include explanatory labels for each parking area. These explanatory labels include the dominant modality source, score stability level, and uncertainty level, for use by the driving system or human review. This module achieves unified decision output from multi-source information and is a crucial supporting link in the entire parking area identification method, moving from identification to recommendation.
[0110] The above scheme maps semantic features and temporal features to a preset space to obtain standard semantic features and standard temporal features. Weighted fusion processing of the standard semantic features and standard temporal features yields a region representation for each physical region, ensuring consistency in dimensionality, scale, and confidence level between the standard semantic features and standard temporal features. The semantic score of the standard semantic feature and the temporal score of the standard temporal feature in the region representation are determined. A first score deviation between the semantic score and the security score is determined, and a second score deviation between the temporal score and the security score is determined. Based on the first and second score deviations, it is determined whether to adjust the preset semantic weights and preset temporal weights to achieve reverse verification of each score and avoid significant score deviations.
[0111] In some embodiments, step 105 includes: Step 105A: Obtain the user's parking preference data, and determine the user preference score and preference weight based on the parking preference data.
[0112] In practical implementation, to further enhance flexibility and user acceptance during actual driving, a personalized preference injection mechanism is introduced into the model decision-making process. This ensures that recommended parking areas not only meet environmental safety constraints but also align as closely as possible with users' current or long-term behavioral preferences. User preference data is explicitly embedded into the data fusion process. Through semantic graph construction and scoring weight adjustment, personalized guidance is added to the target score, thereby achieving truly controllable and interpretable parking area recommendation outputs.
[0113] User preference data includes immediate preferences and historical behavioral preferences. Immediate preferences are specific needs reflected in voice commands or navigation settings; for example, immediate preferences might be voice commands like "approach the rest area," "avoid main roads," or "prioritize areas with cover." Historical behavioral preferences include the characteristics of parking areas chosen by the user in similar environments, records of proactive adjustments to parking suggestions, and responses to system recommendations.
[0114] Semantic modeling of immediate and historical behavioral preferences generates a unified text embedding representation. The text processing uses a large-scale language model trained with visual and linguistic alignment to encode natural language content into fixed-length vectors with contextual understanding and command-oriented capabilities.
[0115] During the semantic graph construction phase, user preference data, acting as special semantic nodes, establishes semantic edge connections with all other nodes in the semantic graph. Edge weights are determined based on the embedding similarity between the user preference data and the semantic representation of each node. The introduction of edges allows user intent to propagate in a structured manner within the semantic graph, directly influencing information flow paths and attention distribution patterns. When aggregating node features, cross-modal attention networks prioritize physical regions closely connected to user preference nodes, thus guiding the semantic graph to focus on user-preferred regions during the encoding phase.
[0116] In the target score determination stage, the score weights are further reconstructed based on preferences. Specifically, a preference matching index is introduced when determining the target score, which calculates the degree of matching between the semantic vector of each candidate node and the user's preference vector. The matching degree serves as a dynamic adjustment factor for the fusion score. If a physical area is physically safe but does not match the user's current preferences (e.g., too close to a noise source or lacking rest facilities), the target score for that physical area will be lowered accordingly; conversely, it will be raised. This adjustment mechanism achieves flexible intervention based on individual needs without changing the original model's judgment logic, making the recommended parking areas more acceptable.
[0117] Step 105B: In response to determining that the security score is less than a preset score threshold, a weighted average is performed on the semantic features, the temporal features, and the security score based on preset semantic weights, preset temporal weights, and preset security weights to obtain the target score for each physical region.
[0118] Step 105C: In response to determining that the security score is greater than or equal to a preset scoring threshold, a weighted average is performed on the semantic features, the temporal features, the security score, and the user preference score based on preset semantic weights, preset temporal weights, preset security weights, and the preference weights to obtain a target score for each physical region.
[0119] In practice, to prevent user preferences from interfering with the basic safety assessment of parking areas, the system implements a safety threshold mechanism. Specifically, when the safety score of a physical area is lower than a preset threshold, the corresponding physical area will not be included in the recommendation range, regardless of how user preferences match. Furthermore, when faced with multiple coexisting or conflicting preferences, the system can introduce a weighted reconciliation strategy. By referencing historical selection frequencies and current instruction priorities, different preference weights are assigned to generate a comprehensive intent embedding, avoiding deviations from the target caused by a single preference.
[0120] The system also supports user preference learning and automatic adaptation mechanisms. During long-term operation, it can continuously record user acceptance of recommended parking areas and update the personalized preference model based on interaction events. Once a sufficient amount of data has been accumulated, a user profile model can be formed, automatically loaded each time the system starts, and the recommendation logic can be fine-tuned in advance to prevent user input each time. This significantly improves automation and user engagement, enabling intelligent safe zone recognition to truly achieve personalized performance.
[0121] Ultimately, the recommendations adjusted by the preference injection mechanism not only possess environmental adaptability and temporal consistency but also take into account individual preferences. While ensuring safety, this enhances the smoothness of human-computer interaction and the transparency of decision-making, providing more collaboratively valuable recommendation information for autonomous driving systems.
[0122] The above scheme acquires user parking preference data and determines user preference scores and weights based on this data. When the safety score is less than a preset score threshold, a weighted average of semantic features, temporal features, and the safety score is calculated based on preset semantic weights, preset temporal weights, and preset safety weights to obtain a target score for each physical area. This ensures that user preference data is no longer considered when the safety score of a physical area does not meet safety standards, thus guaranteeing that recommended parking areas meet safety standards. When the safety score is greater than or equal to the preset score threshold, a weighted average of semantic features, temporal features, the safety score, and the user preference score is calculated based on preset semantic weights, preset temporal weights, preset safety weights, and preference weights to obtain a target score for each physical area.
[0123] In some embodiments, the fusion model is deployed on the vehicle side, and real-time responsiveness on embedded devices is ensured through sparse graphs and the GAT lightweight strategy.
[0124] After achieving model training and personalized preference fusion, the entire recognition and recommendation process needs to be deployed to the in-vehicle system to support real-time inference and continuous operation. Considering the limited computing resources, power stability, and communication bandwidth in the vehicle environment, the model is engineered and optimized, including structural pruning, parameter quantization, inference module separation, and parallel scheduling mechanism design, to ensure that the intelligent safety zone recognition system can operate with low latency and high reliability on the embedded computing platform.
[0125] Before model deployment, the entire system's inference process is divided into three independent sub-modules: a semantic graph construction module, a multimodal fusion inference module, and a region scoring and recommendation module. The semantic graph construction module generates a semantic graph from each frame of multimodal environment perception data. Its operating frequency needs to match the sensor sampling frequency, typically set to 10Hz to 20Hz. This module is a non-neural network structure, primarily consisting of geometric computation and embedding transformations. It has low operating costs and can run on a microcontroller unit (MCU) or a low-power central processing unit (CPU).
[0126] The multimodal fusion inference module is the most computationally intensive part of the system, including a cross-modal attention network, a spatiotemporal graph generation model, and a sparse Gaussian process regression model. To adapt to in-vehicle deployment, a lightweight graph neural network structure replaces part of the original graph attention network (GAT), specifically by reducing the number of attention heads, compressing hidden layer dimensions, and using sparse adjacency matrices to reduce computational paths. In the temporal modeling part, the Transformer is replaced with a hierarchical temporal aggregation network, significantly reducing the number of parameters and computational overhead while maintaining the ability to capture key states. To address the deployment issue of the sparse Gaussian process regression model, a sparse representation method is used to pre-build the regression basis set, activating only a small subset of samples relevant to the current scene in each inference, effectively compressing model loading and runtime resources.
[0127] All neural network modules are quantized before deployment. A post-training quantization strategy compresses floating-point parameters into 8-bit integer format (INT8 format), significantly reducing memory usage and computational latency while maintaining essentially the same precision. The quantized model is then converted to a deployment format compatible with either a Graphics Processing Unit (GPU) or a Neural Network Processing Unit (NPU), ensuring low latency and high throughput when running on the main control platform or an Artificial Intelligence (AI) accelerator.
[0128] The inference module operates using a separate, asynchronous mechanism. Graph construction and perception decoding are executed periodically by the main control unit, while the fusion inference module is processed asynchronously by the AI accelerator. The scoring and recommendation modules run as lightweight tasks in the main decision-making thread, sharing scheduling resources with the vehicle control module. The entire process achieves data flow between modules through shared memory and event notification mechanisms, avoiding system jitter caused by frequent I / O operations. Under high-load scenarios, the system can automatically switch to a reduced-frequency mode, performing inference only on high-confidence candidate regions to ensure stable operation of core functions.
[0129] During operation, the system continuously monitors key performance indicators, including average frame processing time, latency fluctuation range, model output stability, and resource utilization. If an anomaly is detected, a rapid rollback mechanism is triggered, downgrading to rule-based regional recommendation logic to prevent model failure in extreme cases from affecting driving decisions. Simultaneously, the system provides an Over-the-Air (OTA) upgrade interface, allowing the backend to continuously optimize and iterate the model and push update packages to the vehicle for automatic deployment, ensuring continuous algorithm evolution and adaptability to real-world conditions.
[0130] The final deployed model processes an average of less than 80 milliseconds per frame of data, with single-frame resource consumption not exceeding the preset storage space, meeting the operational requirements of mass-production-level automotive computing platforms. Inference results are output in a structured interface format, including a location description, safety score, uncertainty estimate, and semantic tags for each candidate region. These are available for use by the autonomous driving control module or human-machine interface system to achieve final safe zone recommendation and action decision-making. The entire deployment scheme balances real-time performance, stability, and maintainability, providing comprehensive support for the system's implementation in real-world road environments.
[0131] Through the above embodiments, environmental data surrounding the vehicle and the vehicle's own state data are acquired, and a semantic graph is constructed based on the environmental and state data. Using a pre-trained cross-modal attention network, the node features corresponding to multiple nodes in the semantic graph are fused to obtain the semantic features of each physical region. This effectively integrates multi-source environmental and state data, enhancing the understanding of complex environments. Using a pre-trained spatiotemporal graph generation model, the temporal features of each physical region are determined based on the semantic features, capturing the dynamic changes of physical regions at different times and providing a temporal dimension for safety assessment. Using a pre-trained safety scoring model, the safety score of each physical region is determined based on the temporal features, enabling a quantitative assessment of the safety status of physical regions. A target score for each physical region is determined based on semantic features, temporal features, and safety scores. Parking areas are then determined from multiple physical regions based on the target scores. This multi-dimensional feature-based approach makes the target scores more accurate, significantly improving the accuracy and rationality of parking area selection, effectively reducing safety risks during parking, and enhancing the efficiency and safety of vehicle parking.
[0132] It should be noted that the embodiments of this disclosure can also be further described in the following ways: Figure 2 This is a flowchart illustrating a parking area identification method based on multimodal data, as described in an embodiment of this disclosure. Figure 2 As shown, the parking area identification method based on multimodal data includes: Step 1: Multi-source sensing data acquisition and standardization processing.
[0133] Image data, point cloud data, map data, pose data, and user preference data are collected from vehicle-mounted sensors, and spatiotemporal alignment is unified to standardize modal features.
[0134] Step 2: Construction of the multimodal semantic graph structure.
[0135] Construct a semantic graph structure composed of nodes of various modalities such as images, text, point clouds, and maps, and define node types and edge weight rules.
[0136] Step 3, training the cross-modal graph attention encoder.
[0137] A cross-modal attention network is used to extract multimodal attention features from the semantic graph, aggregating semantic and spatial information from different modalities.
[0138] Step 4: Spatiotemporal graph Transformer dynamic scene modeling.
[0139] A spatiotemporal graph generation model is introduced to model the semantic graph of consecutive frames, and the semantic evolution of the region state over time is learned.
[0140] Step 5: Regional safety score label generation and annotation mechanism, sparse Gaussian process regression modeling for safety score prediction.
[0141] Based on historical parking safety data and collision risk records from vehicle manufacturers, the safety scores of parking areas in a scenario are automatically labeled as supervision targets. A sparse Gaussian process regression model is used to model each parking area in the embedding space.
[0142] Step 6: Multi-model fusion mechanism and reverse consistency optimization.
[0143] By fusing the outputs of three models—a cross-modal attention network, a spatiotemporal graph generation model, and a sparse Gaussian process regression model—a reverse residual optimization module is designed to improve recommendation confidence and consistency.
[0144] Step 7: Personalized user preference injection and weight reconstruction mechanism.
[0145] By embedding user preference text into the initialization and post-fusion stages of the semantic graph, the graph structure and model reasoning process are reconstructed to achieve personalized recommendations.
[0146] Step 8: Deployment of the online inference module and model quantization.
[0147] The fusion model is deployed on the vehicle side, and real-time response capability on embedded devices is ensured through sparse graphs and GAT lightweight strategies.
[0148] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0149] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0150] Based on the same inventive concept, corresponding to any of the above-described embodiments, this disclosure also provides a parking control device.
[0151] refer to Figure 3The parking control device includes: The semantic graph construction module 301 is configured to acquire environmental data around the vehicle and the vehicle's own state data, and construct a semantic graph based on the environmental data and the state data; wherein, the semantic graph includes multiple physical regions; The semantic feature determination module 302 is configured to use a pre-trained cross-modal attention network to fuse the node features corresponding to multiple nodes in the semantic graph to obtain the semantic features of each physical region. The temporal feature determination module 303 is configured to use a pre-trained spatiotemporal graph generation model to determine the temporal features of each physical region based on the semantic features. The security score determination module 304 is configured to determine the security score of each physical region based on the temporal features using a pre-trained security score model. The target score determination module 305 is configured to determine the target score of each physical region based on the semantic features, the temporal features and the security score, and to determine the parking area from multiple physical regions based on the target score.
[0152] In some embodiments, the semantic graph construction module 301 includes: The data acquisition unit is configured to acquire environmental data around the vehicle and the vehicle's own status data; wherein the environmental data includes image data and point cloud data, and the status data includes map data. The standardization processing unit is configured to perform time synchronization processing on the image data, the point cloud data, and the map data to obtain time synchronization data, and to perform standardization processing on the time synchronization data to obtain standard image data, standard point cloud data, and standard map data. The semantic graph construction unit is configured to construct a semantic graph based on the standard image data, the standard point cloud data, and the standard map data.
[0153] In some embodiments, the semantic graph construction unit includes: The image node determination subunit is configured to encode the standard image data to obtain an image embedding vector, and use the image embedding vector as an image node; The text node determination subunit is configured to extract text embedding vectors from the standard map data and use the text embedding vectors as text nodes. The physical node determination subunit is configured to determine the point cloud cluster center and cluster feature vector based on the standard point cloud data, and determine the physical region based on the point cloud cluster center and the cluster feature vector, and use the physical region as a physical node; The map node determination subunit is configured to extract structured labels from the standard map data and map the structured labels to map nodes; The edge determination subunit is configured to determine the spatial similarity, semantic relevance, and temporal similarity between every two nodes, and to determine the edge between every two nodes based on the spatial similarity, the semantic relevance, and the temporal similarity. The semantic graph construction subunit is configured to construct a semantic graph based on the image nodes, the text nodes, the physical nodes, the map nodes, and the edges between every two nodes.
[0154] In some embodiments, the semantic feature determination module 302 includes: The feature mapping unit is configured to use a pre-trained cross-modal attention network to map image node features, text node features, physical node features and map node features to a preset semantic space to obtain standard node features; The semantic feature determination unit is configured to determine the neighbor node features of each standard node feature, fuse the standard node features and the neighbor node features to obtain updated node features, and use the updated node features as the semantic features of the physical region corresponding to the standard node features.
[0155] In some embodiments, the timing feature determination module 303 includes: The temporal node feature determination unit is configured to use a pre-trained spatiotemporal graph generation model to determine nodes of the same type from the semantic features of multiple consecutive frames, and to determine the temporal node features of each physical region based on the nodes of the same type. The spatial node feature determination unit is configured to determine the spatial node features of each physical region from the semantic features of multiple consecutive frames using a pre-trained spatiotemporal graph generation model. The timing feature determination unit is configured to determine the timing features of each physical region based on the timing node features and the spatial node features.
[0156] In some embodiments, the security scoring determination module 304 includes: The security score determination unit is configured to use a pre-trained security score model to determine the security score and confidence level of each physical region based on the time-series features. The priority determination unit is configured to sort multiple physical regions according to the security score and the confidence level to obtain the priority of the multiple physical regions.
[0157] In some embodiments, the target scoring determination module 305 includes: The region representation determination unit is configured to map the semantic features and the temporal features to a preset space to obtain standard semantic features and standard temporal features, and to perform weighted fusion processing on the standard semantic features and standard temporal features to obtain a region representation for each physical region; The scoring deviation determination unit is configured to determine the semantic score of the standard semantic feature and the temporal score of the standard temporal feature in the region representation, determine a first scoring deviation between the semantic score and the security score, and determine a second scoring deviation between the temporal score and the security score. The first target score determination unit is configured to determine the target score of each physical region based on a preset semantic weight, a preset temporal weight, and a preset security weight in response to determining that both the first score deviation and the second score deviation are less than a preset score threshold. The second target score determination unit is configured to determine the target score of each physical region based on a first update semantic weight, a first update temporal weight, and a preset security weight in response to determining that the first score deviation is greater than or equal to a preset deviation threshold and the second score deviation is less than a preset score threshold; wherein the first update semantic weight is less than the preset semantic weight and the first update temporal weight is greater than the preset temporal weight. The third target score determination unit is configured to determine the target score of each physical region based on the second updated semantic weight, the second updated temporal weight, and the preset security weight in response to determining that the second score deviation is greater than or equal to a preset deviation threshold and the first score deviation is less than a preset score threshold; wherein the second updated semantic weight is greater than the preset semantic weight and the second updated temporal weight is less than the preset temporal weight.
[0158] In some embodiments, the target scoring determination module 305 includes: The user preference rating determination unit is configured to acquire user parking preference data and determine user preference rating and preference weight based on the parking preference data; The fourth target scoring determination unit is configured to, in response to determining that the security score is less than a preset scoring threshold, perform a weighted average of the semantic features, the temporal features and the security score based on preset semantic weights, preset temporal weights and preset security weights to obtain a target score for each physical region. The fifth target score determination unit is configured to, in response to determining that the security score is greater than or equal to a preset score threshold, perform a weighted average of the semantic features, the temporal features, the security score, and the user preference score based on preset semantic weights, preset temporal weights, preset security weights, and the preference weights to obtain a target score for each physical region.
[0159] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0160] The apparatus described above is used to implement the corresponding parking control method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0161] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the parking control method described in any of the above embodiments.
[0162] Figure 4 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0163] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0164] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0165] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0166] The communication interface 1040 is used to connect the communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0167] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0168] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0169] The electronic devices described above are used to implement the corresponding parking control methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0170] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the parking control method as described in any of the above embodiments.
[0171] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0172] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the parking control method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0173] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a vehicle, including the parking control device, electronic device, or storage medium in the above embodiments, wherein the vehicle device implements the parking control method described in any of the above embodiments.
[0174] The vehicles described in the above embodiments are used to implement the parking control method described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0175] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions. When the computer program instructions are run on a computer, the computer executes the parking control method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0176] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0177] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0178] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0179] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0180] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0181] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0182] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0183] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A parking control method, characterized in that, The method includes: The system acquires environmental data surrounding the vehicle and the vehicle's own state data, and constructs a semantic graph based on the environmental data and the state data; wherein the semantic graph includes multiple physical regions; Using a pre-trained cross-modal attention network, the node features corresponding to multiple nodes in the semantic graph are fused to obtain the semantic features of each physical region; Using a pre-trained spatiotemporal graph generation model, the temporal characteristics of each physical region are determined based on the semantic features; Using a pre-trained safety scoring model, a safety score for each physical region is determined based on the temporal features. Each physical node is assigned a safety score value to characterize the feasibility of the physical region corresponding to the node as a parking area during the current and predicted time periods. A sparse Gaussian process regression model is introduced to predict regional safety in a multimodal embedding space. The input to the sparse Gaussian process regression model is the semantic and temporal features output by the cross-modal attention network and the spatiotemporal graph generation model, and the output is a continuous-valued safety score. A target score for each physical region is determined based on the semantic features, the temporal features, and the security score; and a parking area is determined from multiple physical regions based on the target score. The step of determining the target score for each physical region based on the semantic features, the temporal features, and the security score includes: Obtain user parking preference data, and determine user preference score and preference weight based on the parking preference data; In response to determining that the security score is less than a preset score threshold, a weighted average is performed on the semantic features, the temporal features, and the security score based on preset semantic weights, preset temporal weights, and preset security weights to obtain a target score for each physical region. In response to determining that the security score is greater than or equal to a preset scoring threshold, a weighted average is performed on the semantic features, the temporal features, the security score, and the user preference score based on preset semantic weights, preset temporal weights, preset security weights, and the preference weights to obtain a target score for each physical region.
2. The method according to claim 1, characterized in that, The process of acquiring environmental data surrounding the vehicle and the vehicle's own state data, and constructing a semantic graph based on the environmental data and the state data, includes: Acquire environmental data surrounding the vehicle and the vehicle's own status data; wherein, the environmental data includes image data and point cloud data, and the status data includes map data; The image data, point cloud data, and map data are time-synchronized to obtain time-synchronized data, and the time-synchronized data is standardized to obtain standard image data, standard point cloud data, and standard map data. A semantic graph is constructed based on the standard image data, the standard point cloud data, and the standard map data.
3. The method according to claim 2, characterized in that, The construction of a semantic graph based on the standard image data, the standard point cloud data, and the standard map data includes: The standard image data is encoded to obtain an image embedding vector, and the image embedding vector is used as an image node; Extract text embedding vectors from the standard map data and use the text embedding vectors as text nodes; Based on the standard point cloud data, the point cloud cluster centers and cluster feature vectors are determined, and physical regions are determined according to the point cloud cluster centers and cluster feature vectors, and the physical regions are used as physical nodes; Structured labels are extracted from the standard map data, and the structured labels are mapped to map nodes; Determine the spatial similarity, semantic relevance, and temporal similarity between any two nodes, and determine the edges between any two nodes based on the spatial similarity, semantic relevance, and temporal similarity. A semantic graph is constructed based on the image nodes, text nodes, physical nodes, map nodes, and the edges between every two nodes.
4. The method according to claim 1, characterized in that, The method of using a pre-trained cross-modal attention network to fuse node features corresponding to multiple nodes in the semantic graph to obtain semantic features for each physical region includes: A pre-trained cross-modal attention network is used to map image node features, text node features, physical node features, and map node features to a predefined semantic space to obtain standard node features; The neighboring node features of each standard node feature are determined, and the standard node features and the neighboring node features are fused to obtain the updated node features. The updated node features are then used as the semantic features of the physical region corresponding to the standard node features.
5. The method according to claim 1, characterized in that, The method of using a pre-trained spatiotemporal graph generation model to determine the temporal features of each physical region based on the semantic features includes: Using a pre-trained spatiotemporal graph generation model, nodes of the same type are identified from the semantic features of multiple consecutive frames, and temporal node features of each physical region are determined based on the nodes of the same type. Using a pre-trained spatiotemporal graph generation model, spatial node features of each physical region are determined from the semantic features of multiple consecutive frames; The temporal characteristics of each physical region are determined based on the temporal node characteristics and the spatial node characteristics.
6. The method according to claim 1, characterized in that, The process of determining the security score for each physical region based on the temporal features using a pre-trained security scoring model includes: Using a pre-trained security scoring model, the security score and confidence level of each physical region are determined based on the time-series characteristics. The multiple physical regions are sorted according to the security score and the confidence level to obtain the priority of the multiple physical regions.
7. The method according to claim 1, characterized in that, The step of determining the target score for each physical region based on the semantic features, the temporal features, and the security score includes: The semantic features and the temporal features are mapped to a preset space to obtain standard semantic features and standard temporal features. The standard semantic features and the standard temporal features are then weighted and fused to obtain the region representation of each physical region. Determine the semantic score of the standard semantic feature and the temporal score of the standard temporal feature in the region representation, determine the first score deviation between the semantic score and the security score, and determine the second score deviation between the temporal score and the security score; In response to determining that both the first scoring deviation and the second scoring deviation are less than a preset scoring threshold, a target score for each physical region is determined based on a preset semantic weight, a preset temporal weight, and a preset security weight. In response to determining that the first scoring deviation is greater than or equal to a preset deviation threshold and the second scoring deviation is less than a preset scoring threshold, a target score for each physical region is determined based on a first updated semantic weight, a first updated temporal weight, and a preset security weight; wherein, the first updated semantic weight is less than the preset semantic weight, and the first updated temporal weight is greater than the preset temporal weight. In response to determining that the second scoring deviation is greater than or equal to a preset deviation threshold and the first scoring deviation is less than a preset scoring threshold, a target score for each physical region is determined based on a second updated semantic weight, a second updated temporal weight, and a preset security weight; wherein the second updated semantic weight is greater than the preset semantic weight, and the second updated temporal weight is less than the preset temporal weight.
8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to perform the method according to any one of claims 1 to 7.
9. A vehicle, characterized in that, Includes the storage medium as described in claim 8.
Citation Information
Patent Citations
Vehicle-machine collaborative navigation method, device and equipment based on cross-view space-time modeling
CN118816932A
Robot sensing and decision-making method based on lightweight multi-modal large model
CN120612683A